Pith. sign in

Paper Citation Record · LEDGER

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning

As of 8 August 2026, this Paper Citation Record lists 65 of 65 outbound references and 1 inbound Pith citation observation for arXiv:2507.18531.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.18531 v1

Coverage vector

measured 65 of 65 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T14:36:29.862040Z

measured 66 of 66 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-21T04:35:39.356919Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-21T04:39:35.268570Z

Reference resolution

65 of 65 outbound references displayed

  • verified exact3
  • verified fuzzy4
  • unresolved58
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation dc6a7e31-8201-4180-ba38-9dafbbc109b2 · outbound

This paper cites Qwen2.5-VL Technical Report.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Qwen2.5-VL Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T14:36:28.756643Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:36:28.756643Z digest=sha256:8f5ab887ea530a87c960cd6b3613d1ac7112731430a05e4cf815b327ca729ca2

Observation 42204628-102c-42c1-b5f6-45ebfafbd812 · outbound

This paper cites an unresolved cited work.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Unresolved cited work

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T14:36:28.774002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:36:28.774002Z digest=sha256:5c2653e77c2241f6fc2560f5e301b10fab91df391f0dd5d8977153b658bca6d4

Observation 2f414133-f829-442a-b0ed-51acb9ff197f · outbound

This paper cites an unresolved cited work.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:36:33.630954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:36:28.804732Z digest=sha256:d2993d1020a799f4c6c74b64da756cece5e1b04c653656347578df66da0057e8

Observation bbb85bc4-a452-4b8e-8689-a4428efd383a · outbound

This paper cites VideoLLM: Modeling Video Sequence with Large Language Models.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning VideoLLM: Modeling Video Sequence with Large Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T14:36:28.824818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:36:28.824818Z digest=sha256:b1717514032e501e0cca42667f266ca742f7e8064638f4090b1d7504e12d45a7

Observation 13215b2a-2602-4f18-bf25-eeb3a07d5c9b · outbound

This paper cites an unresolved cited work.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Unresolved cited work

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T14:36:28.830571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:36:28.830571Z digest=sha256:d0ad891544169892567a508b4484afe8a501c8ce195f97243ab00537c0a35988

Observation 731f6ea5-715e-42e1-b475-8aae446a1e72 · outbound

This paper cites an unresolved cited work.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:36:33.598539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:36:28.862647Z digest=sha256:d929941a15e2e34e5123c538d9c2cdb4e90357f5bdec59df73c08fd11b886390

Observation a16528c4-38d0-4fba-ac9f-3c2da59e467c · outbound

This paper cites an unresolved cited work.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:36:33.567933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:36:28.870648Z digest=sha256:7272e9bfdc73232b99a57b2a8531ed4a76539cd97893aba82db915aee039289a

Observation 0c156231-190f-4f73-82ca-c18507dc9569 · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T14:36:28.875907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:36:28.875907Z digest=sha256:4749bb2a4c54baacf30d8eefc564cf4671205ff898cb05341f36d6232f4e8930

Observation 4ad8fe91-e6ad-465f-808c-5951059f9066 · outbound

This paper cites an unresolved cited work.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:36:33.519400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:36:28.884001Z digest=sha256:afcccfbd115a3bc5e0263efea907d70a3179e2164046af65a6de56d9d480fe8f

Observation 76d80ee7-7cae-45e7-a7ff-4b17429fda7c · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T14:36:28.890640Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:36:28.890640Z digest=sha256:92339054ed2463f3179f91715fece4ec730e9797ca2c79182d4184762d9bb9ba

Observation 0523fc10-059e-42a5-8757-946baf492dc1 · outbound

This paper cites an unresolved cited work.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:36:33.484662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:36:28.896730Z digest=sha256:694786be8af90ee66730c52021772ade9f51b93db0f6eda06871361b2d11a85c

Observation 5abe577e-90ef-4cc7-bd88-a015cf6a7550 · outbound

This paper cites an unresolved cited work.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Unresolved cited work

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T14:36:28.902907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:36:28.902907Z digest=sha256:efb7d8224fcd110a4c6b68f8777af008a756a57b90a14970c19f6b370d9a4d1c

Observation 308f1e94-137a-42e2-afa2-859c0e6185df · outbound

This paper cites an unresolved cited work.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Unresolved cited work

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T14:36:28.911252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:36:28.911252Z digest=sha256:02c5f20d4d10aa12493fabaa9731ad3bbd28f64ab88119386815f5963754aeb3

Observation 30c5393b-afda-49a0-8f06-22ab38060af6 · outbound

This paper cites Kastner, Yasutomo Kawanishi, Trung Thanh Nguyen, and Junan Chen.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Kastner, Yasutomo Kawanishi, Trung Thanh Nguyen, and Junan Chen

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:36:33.409054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:36:28.918071Z digest=sha256:5d1121ed571c7ad5105e3650c63588b574b7c5b2dc80b4d74bdd2ffc3feff8aa

Observation b4dad6c9-5a3b-4527-bb40-5c391699ec5b · outbound

This paper cites an unresolved cited work.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Unresolved cited work

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T14:36:28.935806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:36:28.935806Z digest=sha256:8110471985e1cb26fa59b947eb27144704d8b18d0a2408f4406e3e34827a6a43

Observation dfe291c8-0c5b-44b1-a2eb-1298ca80235e · outbound

This paper cites an unresolved cited work.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:36:33.330172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:36:28.950400Z digest=sha256:af451ba2d1f40a8f9a918c9b37e756189201de5017f05a6c7e986ef0e309ac15

Observation 93078549-2e74-4600-8872-1a6b7649fb2a · outbound

This paper cites an unresolved cited work.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:36:33.136958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:36:28.957212Z digest=sha256:58b4f1a07d9bf0ca4faad6c49053347560982d00efc457be752de58e88b6712c

Observation 288323a4-0bae-499f-a36d-8327cb51eeba · outbound

This paper cites The First Place Solution of WSDM Cup 2024: Leveraging Large Language Models for Conversational Multi-Doc QA.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning The First Place Solution of WSDM Cup 2024: Leveraging Large Language Models for Conversational Multi-Doc QA

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-08-06T14:36:30.683604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:36:28.963157Z digest=sha256:2f0011e860bed0e0711b86565df6f89cd020bb9f805a890e9671cde38b24fa63

Observation 6149dfaf-44e7-4ebe-91ac-a649344baf6b · outbound

This paper cites GroundingGPT:Language Enhanced Multi-modal Grounding Model.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning GroundingGPT:Language Enhanced Multi-modal Grounding Model

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T14:36:28.980544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:36:28.980544Z digest=sha256:b99026cce235b47006b84c48e6b6a327d56c37364f47c235b4dd91c9286fd515

Observation 6e131865-637e-4ff8-b0d5-85303bf7a887 · outbound

This paper cites Video-LLaVA: Learning United Visual Representation by Alignment Before Projection.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T14:36:28.987835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:36:28.987835Z digest=sha256:49430ac1e6bba154a2cab3501a97a1739e80ec0fa24923e3aa717e2e66998f51

Observation 8577305e-dd05-4d2f-afa0-f5367d9c21b6 · outbound

This paper cites an unresolved cited work.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:36:32.907929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:36:29.038967Z digest=sha256:cbb0171fd1ba3bf357014d05579f0af62c7a22e516be91a8f59cad53bd455339

Observation 0bd1384c-21e0-418e-a515-4dca5363be3d · outbound

This paper cites an unresolved cited work.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Unresolved cited work

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T14:36:29.074740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:36:29.074740Z digest=sha256:6a0bb85d959dbcd02b0a2808fefb4bcf99465841f5b1b69503ae8a3018eae496

Observation 9a5d55e2-8499-4997-8b8f-7e832017e2e5 · outbound

This paper cites Macaw-LLM: Multi-Modal Language Modeling with Image, Audio, Video, and Text Integration.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Macaw-LLM: Multi-Modal Language Modeling with Image, Audio, Video, and Text Integration

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T14:36:29.101479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:36:29.101479Z digest=sha256:7e0340127550fa67c0af6444215516289a727f37b1ee3ddecbd990bf8584f166

Observation cfb91417-6894-4509-bd7d-a56b5435c770 · outbound

This paper cites an unresolved cited work.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Unresolved cited work

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T14:36:29.124744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:36:29.124744Z digest=sha256:e2adbdfe5903efdd569d1a0ae7449ba53e81301213777463dda93f9396322672

Observation ae3ba893-d498-469e-b526-bc5454769ffc · outbound

This paper cites an unresolved cited work.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Unresolved cited work

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T14:36:29.149139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:36:29.149139Z digest=sha256:446c67dc4941f8afd2494c5e67895ecfaca689686d516ed399994d750ab84802

Observation 5e52972a-9ce2-4cb0-94fb-a9ebb190d69b · outbound

This paper cites Reinforced Video Captioning with Entailment Rewards.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Reinforced Video Captioning with Entailment Rewards

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-08-06T14:36:30.540454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:36:29.160509Z digest=sha256:7caf45bbbfd28f71e0356a87a68cfce3eb5f4204d89b2bfbf2bcb36ca142160a

Observation bac88a13-b8bd-4d48-a736-82ca1ff6b729 · outbound

This paper cites Kosmos-2: Grounding Multimodal Large Language Models to the World.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Kosmos-2: Grounding Multimodal Large Language Models to the World

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T14:36:29.179246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:36:29.179246Z digest=sha256:c65bd487ca2ffa35040fdd8f328b33606ae126480e5c16f477e95f19917132bd

Observation 0481b339-1842-4dcd-a94c-bd18a6cd7b7a · outbound

This paper cites an unresolved cited work.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:36:32.645965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:36:29.204885Z digest=sha256:e173c4647e8c9e9f08b21f150324c8d10411923ae782a627040ef14c6c936d38

Observation bc1ed0eb-b562-42ae-a323-31a5a96752f1 · outbound

This paper cites Audio-Visual LLM for Video Understanding.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Audio-Visual LLM for Video Understanding

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T14:36:29.216841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:36:29.216841Z digest=sha256:f3aec34b20b9a6cc93ea5d72528baf9edcaab3a6d3d6aa0d410beff7dad3d606

Observation f0a0e0de-24b5-4f19-a70d-d297d9ff9c47 · outbound

This paper cites an unresolved cited work.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Unresolved cited work

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T14:36:29.236144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:36:29.236144Z digest=sha256:33f36badc711c2d5e9ecb8f6d70c34b4cacfb4e78fc835e257324a389ca3c192

Observation 4374199a-da92-4eb0-95ac-ec8cef97716d · outbound

This paper cites an unresolved cited work.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Unresolved cited work

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T14:36:29.279165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:36:29.279165Z digest=sha256:4f5e4d90e6c031d65634b0ea98c94dc88bba86febceecaff74c31772bee00f2a

Observation 3dc758d5-3d19-4cd6-847a-be6cc3d62e1e · outbound

This paper cites an unresolved cited work.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Unresolved cited work

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T14:36:29.296001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:36:29.296001Z digest=sha256:e6382e2b280868233a699f93b23da2bf7135c8b47d21c9b8d8f7ce812fd5468a

Observation 07d0839b-d37b-44fd-9e28-5f60c3c7f7dd · outbound

This paper cites an unresolved cited work.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Unresolved cited work

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T14:36:29.332967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:36:29.332967Z digest=sha256:56b47e5db505574371641df10e5e2e2cb3539e0782225b35f5048398745a89f2

Observation 06e59a2a-f5cb-47b1-8749-73b6fb7f2721 · outbound

This paper cites In Proceedings of the 31st ACM International Conference on Multimedia.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning In Proceedings of the 31st ACM International Conference on Multimedia

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T14:36:29.326969Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:36:29.326969Z digest=sha256:0f700c4ec3b8f91825eb521f372309b78f98503674c56b66cea425aabe4d54b1

Observation ad3ad4e7-b1ae-4c7d-b6bc-9f6e5a73b46a · outbound

This paper cites Caption Anything in Video: Fine-grained Object-centric Captioning via Spatiotemporal Multimodal Prompting.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Caption Anything in Video: Fine-grained Object-centric Captioning via Spatiotemporal Multimodal Prompting

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T14:36:29.394779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:36:29.394779Z digest=sha256:1622238aa724befcefd8af782a45c7433352c12929e2d15a4f60e837df33850c

Observation 095751ea-1c5c-419e-8643-040b6bf1c3d9 · outbound

This paper cites an unresolved cited work.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Unresolved cited work

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T14:36:29.447467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:36:29.447467Z digest=sha256:67a7c4528022dc11b1ca0397051ac83e53b6b25e2d80ce7100f7a958f7d512c3

Observation 319a4095-870c-416d-a64c-ca4636aa7a69 · outbound

This paper cites an unresolved cited work.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:36:32.040258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:36:29.365989Z digest=sha256:a1b3eb8db48d3fe0ac1bdc4cb05501819ebbe77774e5b4e71eda7c2d2184a2ad

Observation 31188037-4dab-4cc2-ad2d-635ee4f64788 · outbound

This paper cites an unresolved cited work.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Unresolved cited work

Reference 38

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:36:31.694759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:36:29.491095Z digest=sha256:1679acf6917da80ca124b9dc54b75de451741d0930414af102653428421f0351

Observation f38be541-5772-4fe8-9f6b-4b15eb7bf32b · outbound

This paper cites an unresolved cited work.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:36:31.549968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:36:29.499869Z digest=sha256:eee142d3572fac76a765dab0b62d9fd499f339799955de34825edf0f8cac80b6

Observation 19b5d2ad-b7c1-4368-8a6c-07d279cc7f99 · outbound

This paper cites an unresolved cited work.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Unresolved cited work

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T14:36:29.474796Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:36:29.474796Z digest=sha256:b168dfeb16a7bcf911733e05ba63fbd02df29d9f5812a71cb3cc2e82cfce3d3d

Observation c973e061-80aa-4b03-9b71-d641c42539c6 · outbound

This paper cites an unresolved cited work.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Unresolved cited work

Reference 41

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:36:31.138936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:36:29.532008Z digest=sha256:5da3491b04e2c2b0113bc0b5777671367bca8a47876add715acb98e8639d665a

Observation 9e461aa4-bd38-4e8b-a4bb-e5313aa4cb95 · outbound

This paper cites InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T14:36:29.583224Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:36:29.583224Z digest=sha256:0c5225ad47ec1f00abed70a7781db00e6fbea3760acd473bb8d09d25a299f346

Observation c9b73b3b-3978-44c9-8fe7-ea8da1860696 · outbound

This paper cites an unresolved cited work.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Unresolved cited work

Reference 43

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:36:31.450436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:36:29.509421Z digest=sha256:a3cc19fa64edf23c4c7e4c4827c4ed3454df92f665bcced4afd3c9d3d8672195

Observation c5cf68d7-08d0-45c3-9875-4330fd4e649a · outbound

This paper cites an unresolved cited work.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Unresolved cited work

Reference 44

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:36:31.038003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:36:29.654739Z digest=sha256:f7889eeb7d4b3fd244954c5eb8be1e4314e0d5aaa7e6ae5a3b415b8794403f13

Observation 53ec807f-a90b-4c6d-a069-f22338da95a8 · outbound

This paper cites SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T14:36:29.673490Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:36:29.673490Z digest=sha256:cac34d147e289208e5e889e5d5ae88b8770c448e302df2198cd80adbeb5a5c7a

Observation 68df01c5-db36-4869-a1ba-cc133ae62b8f · outbound

This paper cites CLIP-ViP: Adapting Pre-trained Image-Text Model to Video-Language Representation Alignment.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning CLIP-ViP: Adapting Pre-trained Image-Text Model to Video-Language Representation Alignment

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T14:36:29.684895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:36:29.684895Z digest=sha256:6645822b59cf35559de022b71a644fd68d37f0efb056956ff48ccfd6bda97cb5

Observation 5835d953-a9c2-44f1-84de-92bb33d1d9e7 · outbound

This paper cites an unresolved cited work.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Unresolved cited work

Reference 47

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:36:31.110929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:36:29.607606Z digest=sha256:eb9c57169a557ba0bb9ea640839c82afa3f6ea04dc26a03a0f71c8a1c1ae8410

Observation d7067fad-f68f-460a-a4fe-d845ebd89253 · outbound

This paper cites an unresolved cited work.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Unresolved cited work

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T14:36:29.709137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:36:29.709137Z digest=sha256:4e013104fe6bef70b6d9f55a7c0590a485569681f8dbb6a0be4cc49d27c0d705

Observation 31dcf7d2-1e4f-48d5-88ad-189e61538283 · outbound

This paper cites an unresolved cited work.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Unresolved cited work

Reference 49

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:36:30.976741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:36:29.726957Z digest=sha256:7cd6e5e9a237ddd166811366c3ccaefec37f0c9b9bb9abe53f160b1e0c4551e9

Observation 29c554df-1207-4a9c-80f4-2f8247fbcc36 · outbound

This paper cites an unresolved cited work.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Unresolved cited work

Reference 50

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:36:30.946168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:36:29.741989Z digest=sha256:94db9f7aabf6577e9022fa6b8d40bfb7a86886e1e1f36536559c94728dc91638

Observation 03d931d2-76c4-40a9-b6b1-26f88676df9e · outbound

This paper cites an unresolved cited work.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Unresolved cited work

Reference 51

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:36:30.906135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:36:29.748938Z digest=sha256:e7458f7270696d312572b577f42ce2a01792d22de1595ca85c4d4cb68733d43d

Observation a70c7932-59e7-4ac5-a0af-b4d76ac58e99 · outbound

This paper cites Qwen3 Technical Report.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Qwen3 Technical Report

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T14:36:29.695562Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:36:29.695562Z digest=sha256:555e9ee8a033e563cab98ba1259f8c1335687f3e45e2afb26466e86a2c6ec0f2

Observation fcf9863a-fb78-41c6-92da-21c52c478f7c · outbound

This paper cites VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T14:36:29.767987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:36:29.767987Z digest=sha256:f5337387d882975bd15282c6e2ca776511b90f095997bf9c41d3af1aa4fc5b7c

Observation be0ef333-e251-4713-8cf3-b698da8e4b71 · outbound

This paper cites Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T14:36:29.714217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:36:29.714217Z digest=sha256:1d1e8ef37d47a0cbc5ef3abd5ee27a5f07d0c3c3f08848ac0fe4715e2dd765ef

Observation 918efbcc-c026-423b-8c60-ec56077fd46b · outbound

This paper cites an unresolved cited work.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Unresolved cited work

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T14:36:29.787485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:36:29.787485Z digest=sha256:1ca479d1ebd11df371791969215d31e52c9255397d21d03f172dcf2e3b64901e

Observation 5e30fd84-e35e-4b66-8987-d6987b31df39 · outbound

This paper cites an unresolved cited work.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Unresolved cited work

Reference 56

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:36:30.825970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:36:29.797201Z digest=sha256:6cbc17db35f8b38572ebb3ad6ddfcc989064ad634e30381a549badf3feb17f04

Observation 87f6d969-dcfa-47c6-a344-a8830c003976 · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-06T14:36:29.809595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:36:29.809595Z digest=sha256:76e1503b987a2df409bb6ab3874047fa86fdf6a3d359acfe5b54e777ba619d59

Observation 535d5528-c211-40a6-89a1-41072bdde2d6 · outbound

This paper cites Ferret: Refer and Ground Anything Anywhere at Any Granularity.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Ferret: Refer and Ground Anything Anywhere at Any Granularity

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-06T14:36:29.762895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:36:29.762895Z digest=sha256:45abf340821fa469347134089e17fcaa85139b5ed96019fa2b86ab0eacd1cf28

Observation 97411bd4-bb74-4ae5-b95c-028126faebd9 · outbound

This paper cites InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-06T14:36:29.862040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:36:29.862040Z digest=sha256:79342e04995241e9ce50f21f1dfe6559280e47857bdac5d5e8f1aa07896875ab

Observation 03c16232-4046-467a-8847-08fb25ec7307 · outbound

This paper cites an unresolved cited work.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Unresolved cited work

Reference 60

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:36:30.864740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:36:29.780390Z digest=sha256:cef9e461ed2521fb1c4fbb80a6f0bd844730004452adec0729c244143067319c

Observation 3d196621-58ea-4326-ad9c-5b540ec82e8b · outbound

This paper cites OVC-Net: Object-Oriented Video Captioning with Temporal Graph and Detail Enhancement.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning OVC-Net: Object-Oriented Video Captioning with Temporal Graph and Detail Enhancement

Reference 64

Resolution
verified exact
local_arxiv, observed 2026-08-06T14:36:30.025967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:36:29.835315Z digest=sha256:33c1296f05de7b0c5c05c2391579e6f0323e0f2289220e25ec86a71c265c1213

Observation 1c897235-9b2f-48d4-afe8-7c2a0937e9fc · outbound

This paper cites In Proceedings of the IEEE conference on computer vision and pattern recognition.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning In Proceedings of the IEEE conference on computer vision and pattern recognition

Reference 2018

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:36:31.241487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:36:29.515564Z digest=sha256:96006cec441e639c5bd6bcb7f827e6a4c726c62e374501a6ba537f436e871c4f

Observation 9e3a30e2-4a94-4ba2-abf5-c0628384966f · outbound

This paper cites In Proceedings of the IEEE/CVF international conference on computer vision.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning In Proceedings of the IEEE/CVF international conference on computer vision

Reference 2019

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:36:32.304827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:36:29.356331Z digest=sha256:95a45ec067eb0f27532893ce10ded48ae40b00ff922977f7bce944cec7c4e02e

Observation 4d657601-617a-4e66-80b6-9ff221e4d053 · outbound

This paper cites In 2020 IEEE International Conference on Multimedia and Expo (ICME).

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning In 2020 IEEE International Conference on Multimedia and Expo (ICME)

Reference 2020

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:36:31.068892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:36:29.625877Z digest=sha256:8d3c3721333a80280efb463eda1f1583dca3dd9a13b8adbf7ac01efdb13c3949

Observation a3898842-383a-46dc-bcdc-848142eec563 · outbound

This paper cites Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-06T14:36:28.852339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:36:28.852339Z digest=sha256:a0349701c78a077896a5db73e7d7b4c6ed8613499c5109da8ce3a54f201c5b83

Pith citing papers

Observation dc631b3f-aa4e-4c98-9e70-26fcb13e0df4 · inbound

RoadTones: Tone Controllable Text Generation from Road Event Videos cites this paper.

RoadTones: Tone Controllable Text Generation from Road Event Videos IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-21T04:39:35.270214Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-21T04:35:39.356919Z digest=sha256:1fdf0a16682c327a875770897d16a0b3d6dc86ee42b9a5b36a641fa5b4c3cae5