Pith. sign in

Paper Citation Record · LEDGER

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning

As of 10 August 2026, this Paper Citation Record lists 65 of 65 outbound references and 1 inbound Pith citation observation for arXiv:2507.18531.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.18531 v1

Coverage vector

measured 65 of 65 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T14:36:29.862040Z

measured 66 of 66 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-21T04:35:39.356919Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-21T04:39:35.268570Z

Reference resolution

65 of 65 outbound references displayed

  • verified exact3
  • verified fuzzy4
  • unresolved58
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation dc6a7e31-8201-4180-ba38-9dafbbc109b2 · outbound

This paper cites Qwen2.5-VL Technical Report.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Qwen2.5-VL Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T14:36:28.756643Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:36:28.756643Z digest=sha256:8f5ab887ea530a87c960cd6b3613d1ac7112731430a05e4cf815b327ca729ca2

Observation 42204628-102c-42c1-b5f6-45ebfafbd812 · outbound

This paper cites an unresolved cited work.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Unresolved cited work

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T14:36:28.774002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:36:28.774002Z digest=sha256:5c2653e77c2241f6fc2560f5e301b10fab91df391f0dd5d8977153b658bca6d4

Observation 2f414133-f829-442a-b0ed-51acb9ff197f · outbound

This paper cites an unresolved cited work.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:36:33.630954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T14:36:28.804732Z digest=sha256:cd2fc6c26e38f17b2ff5ea831de2b6f01e6f5c6b6d6876bfeb083c6d30a91d09

Observation bbb85bc4-a452-4b8e-8689-a4428efd383a · outbound

This paper cites VideoLLM: Modeling Video Sequence with Large Language Models.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning VideoLLM: Modeling Video Sequence with Large Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T14:36:28.824818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:36:28.824818Z digest=sha256:8a33915d2b9d4df30468e9e67039079975386eae4abcec53bfc76e3faf4ab82b

Observation 13215b2a-2602-4f18-bf25-eeb3a07d5c9b · outbound

This paper cites an unresolved cited work.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Unresolved cited work

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T14:36:28.830571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:36:28.830571Z digest=sha256:d0ad891544169892567a508b4484afe8a501c8ce195f97243ab00537c0a35988

Observation 731f6ea5-715e-42e1-b475-8aae446a1e72 · outbound

This paper cites an unresolved cited work.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:36:33.598539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T14:36:28.862647Z digest=sha256:5a0c0af34101fd03a9537869fa9122cf7eaa4e62820b38969604b02a8f63e447

Observation a16528c4-38d0-4fba-ac9f-3c2da59e467c · outbound

This paper cites an unresolved cited work.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:36:33.567933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T14:36:28.870648Z digest=sha256:2d0f1c760ebba4e162ab3188826231c5666a6ea7d7020da11e460ef9eed5e6ba

Observation 0c156231-190f-4f73-82ca-c18507dc9569 · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T14:36:28.875907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:36:28.875907Z digest=sha256:4749bb2a4c54baacf30d8eefc564cf4671205ff898cb05341f36d6232f4e8930

Observation 4ad8fe91-e6ad-465f-808c-5951059f9066 · outbound

This paper cites an unresolved cited work.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:36:33.519400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T14:36:28.884001Z digest=sha256:c91edc8facd3612f891c55c8685b7e1e49ca22baa14e25ded01669957b4637ff

Observation 76d80ee7-7cae-45e7-a7ff-4b17429fda7c · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T14:36:28.890640Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:36:28.890640Z digest=sha256:a7d78825504eb5b6db45e7193ad3feb3a75a54fb51d5081b6db955c86cee669d

Observation 0523fc10-059e-42a5-8757-946baf492dc1 · outbound

This paper cites an unresolved cited work.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:36:33.484662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T14:36:28.896730Z digest=sha256:93c68ca61541e265138effd8a6c085b965be1b6fc0fbdea3225b70377c7b1488

Observation 5abe577e-90ef-4cc7-bd88-a015cf6a7550 · outbound

This paper cites an unresolved cited work.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Unresolved cited work

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T14:36:28.902907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:36:28.902907Z digest=sha256:efb7d8224fcd110a4c6b68f8777af008a756a57b90a14970c19f6b370d9a4d1c

Observation 308f1e94-137a-42e2-afa2-859c0e6185df · outbound

This paper cites an unresolved cited work.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Unresolved cited work

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T14:36:28.911252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:36:28.911252Z digest=sha256:02c5f20d4d10aa12493fabaa9731ad3bbd28f64ab88119386815f5963754aeb3

Observation 30c5393b-afda-49a0-8f06-22ab38060af6 · outbound

This paper cites Kastner, Yasutomo Kawanishi, Trung Thanh Nguyen, and Junan Chen.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Kastner, Yasutomo Kawanishi, Trung Thanh Nguyen, and Junan Chen

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:36:33.409054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T14:36:28.918071Z digest=sha256:d0f5e3ffb1557646765cf5c5ef2d6a1e7234490f4816b10687318c4ada96e091

Observation b4dad6c9-5a3b-4527-bb40-5c391699ec5b · outbound

This paper cites an unresolved cited work.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Unresolved cited work

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T14:36:28.935806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:36:28.935806Z digest=sha256:8110471985e1cb26fa59b947eb27144704d8b18d0a2408f4406e3e34827a6a43

Observation dfe291c8-0c5b-44b1-a2eb-1298ca80235e · outbound

This paper cites an unresolved cited work.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:36:33.330172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T14:36:28.950400Z digest=sha256:617ecc81989215bb0d33941786a1f01e85fb4e195f240d90c3816f3b7e96e3d4

Observation 93078549-2e74-4600-8872-1a6b7649fb2a · outbound

This paper cites an unresolved cited work.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:36:33.136958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T14:36:28.957212Z digest=sha256:46f9966269527f0861c921e191fc74c6dd9f27bfaca8d76c324c82b239cfcebf

Observation 288323a4-0bae-499f-a36d-8327cb51eeba · outbound

This paper cites The First Place Solution of WSDM Cup 2024: Leveraging Large Language Models for Conversational Multi-Doc QA.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning The First Place Solution of WSDM Cup 2024: Leveraging Large Language Models for Conversational Multi-Doc QA

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-08-06T14:36:30.683604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T14:36:28.963157Z digest=sha256:b88927f93c15a671a2a5a92eb2a8b43acb9e5494ea932dd45e0af4b81eff1d0f

Observation 6149dfaf-44e7-4ebe-91ac-a649344baf6b · outbound

This paper cites GroundingGPT:Language Enhanced Multi-modal Grounding Model.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning GroundingGPT:Language Enhanced Multi-modal Grounding Model

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T14:36:28.980544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:36:28.980544Z digest=sha256:f9a3f34d0da0f3dc0ba06636995f711c029db50f18e817e7646e680449d93538

Observation 6e131865-637e-4ff8-b0d5-85303bf7a887 · outbound

This paper cites Video-LLaVA: Learning United Visual Representation by Alignment Before Projection.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T14:36:28.987835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:36:28.987835Z digest=sha256:49430ac1e6bba154a2cab3501a97a1739e80ec0fa24923e3aa717e2e66998f51

Observation 8577305e-dd05-4d2f-afa0-f5367d9c21b6 · outbound

This paper cites an unresolved cited work.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:36:32.907929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T14:36:29.038967Z digest=sha256:45e18b0f6816004ff3c47af124edaae41f4eb34329e0dda0660c829cb99e94ec

Observation 0bd1384c-21e0-418e-a515-4dca5363be3d · outbound

This paper cites an unresolved cited work.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Unresolved cited work

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T14:36:29.074740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:36:29.074740Z digest=sha256:6a0bb85d959dbcd02b0a2808fefb4bcf99465841f5b1b69503ae8a3018eae496

Observation 9a5d55e2-8499-4997-8b8f-7e832017e2e5 · outbound

This paper cites Macaw-LLM: Multi-Modal Language Modeling with Image, Audio, Video, and Text Integration.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Macaw-LLM: Multi-Modal Language Modeling with Image, Audio, Video, and Text Integration

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T14:36:29.101479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:36:29.101479Z digest=sha256:7e0340127550fa67c0af6444215516289a727f37b1ee3ddecbd990bf8584f166

Observation cfb91417-6894-4509-bd7d-a56b5435c770 · outbound

This paper cites an unresolved cited work.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Unresolved cited work

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T14:36:29.124744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:36:29.124744Z digest=sha256:e2adbdfe5903efdd569d1a0ae7449ba53e81301213777463dda93f9396322672

Observation ae3ba893-d498-469e-b526-bc5454769ffc · outbound

This paper cites an unresolved cited work.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Unresolved cited work

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T14:36:29.149139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:36:29.149139Z digest=sha256:446c67dc4941f8afd2494c5e67895ecfaca689686d516ed399994d750ab84802

Observation 5e52972a-9ce2-4cb0-94fb-a9ebb190d69b · outbound

This paper cites Reinforced Video Captioning with Entailment Rewards.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Reinforced Video Captioning with Entailment Rewards

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-08-06T14:36:30.540454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T14:36:29.160509Z digest=sha256:a6f7e3ad9691cd3fc70ab25c7541ab479da437a42a614b4d7d74c21294bc03f1

Observation bac88a13-b8bd-4d48-a736-82ca1ff6b729 · outbound

This paper cites Kosmos-2: Grounding Multimodal Large Language Models to the World.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Kosmos-2: Grounding Multimodal Large Language Models to the World

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T14:36:29.179246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:36:29.179246Z digest=sha256:c65bd487ca2ffa35040fdd8f328b33606ae126480e5c16f477e95f19917132bd

Observation 0481b339-1842-4dcd-a94c-bd18a6cd7b7a · outbound

This paper cites an unresolved cited work.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:36:32.645965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T14:36:29.204885Z digest=sha256:04d77feebc94f9b2c6a1ccce39c3cf9f0a6b4ca8c381cb9dc89d5de098b7c46c

Observation bc1ed0eb-b562-42ae-a323-31a5a96752f1 · outbound

This paper cites Audio-Visual LLM for Video Understanding.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Audio-Visual LLM for Video Understanding

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T14:36:29.216841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:36:29.216841Z digest=sha256:f3aec34b20b9a6cc93ea5d72528baf9edcaab3a6d3d6aa0d410beff7dad3d606

Observation f0a0e0de-24b5-4f19-a70d-d297d9ff9c47 · outbound

This paper cites an unresolved cited work.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Unresolved cited work

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T14:36:29.236144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:36:29.236144Z digest=sha256:33f36badc711c2d5e9ecb8f6d70c34b4cacfb4e78fc835e257324a389ca3c192

Observation 4374199a-da92-4eb0-95ac-ec8cef97716d · outbound

This paper cites an unresolved cited work.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Unresolved cited work

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T14:36:29.279165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:36:29.279165Z digest=sha256:4f5e4d90e6c031d65634b0ea98c94dc88bba86febceecaff74c31772bee00f2a

Observation 3dc758d5-3d19-4cd6-847a-be6cc3d62e1e · outbound

This paper cites an unresolved cited work.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Unresolved cited work

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T14:36:29.296001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:36:29.296001Z digest=sha256:e6382e2b280868233a699f93b23da2bf7135c8b47d21c9b8d8f7ce812fd5468a

Observation 07d0839b-d37b-44fd-9e28-5f60c3c7f7dd · outbound

This paper cites an unresolved cited work.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Unresolved cited work

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T14:36:29.332967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:36:29.332967Z digest=sha256:56b47e5db505574371641df10e5e2e2cb3539e0782225b35f5048398745a89f2

Observation 06e59a2a-f5cb-47b1-8749-73b6fb7f2721 · outbound

This paper cites In Proceedings of the 31st ACM International Conference on Multimedia.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning In Proceedings of the 31st ACM International Conference on Multimedia

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T14:36:29.326969Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:36:29.326969Z digest=sha256:0f700c4ec3b8f91825eb521f372309b78f98503674c56b66cea425aabe4d54b1

Observation ad3ad4e7-b1ae-4c7d-b6bc-9f6e5a73b46a · outbound

This paper cites Caption Anything in Video: Fine-grained Object-centric Captioning via Spatiotemporal Multimodal Prompting.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Caption Anything in Video: Fine-grained Object-centric Captioning via Spatiotemporal Multimodal Prompting

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T14:36:29.394779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:36:29.394779Z digest=sha256:1622238aa724befcefd8af782a45c7433352c12929e2d15a4f60e837df33850c

Observation 095751ea-1c5c-419e-8643-040b6bf1c3d9 · outbound

This paper cites an unresolved cited work.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Unresolved cited work

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T14:36:29.447467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:36:29.447467Z digest=sha256:67a7c4528022dc11b1ca0397051ac83e53b6b25e2d80ce7100f7a958f7d512c3

Observation 319a4095-870c-416d-a64c-ca4636aa7a69 · outbound

This paper cites an unresolved cited work.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:36:32.040258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T14:36:29.365989Z digest=sha256:7f3ebbcc5f84ebf3dfd994fa59c4f442906b9d12ecc8f608315460c0550e4965

Observation 31188037-4dab-4cc2-ad2d-635ee4f64788 · outbound

This paper cites an unresolved cited work.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Unresolved cited work

Reference 38

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:36:31.694759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T14:36:29.491095Z digest=sha256:2783ddbd558553218cab286978c1ec7e97eed62763b9c4ff22f8013b8f5dc373

Observation f38be541-5772-4fe8-9f6b-4b15eb7bf32b · outbound

This paper cites an unresolved cited work.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:36:31.549968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T14:36:29.499869Z digest=sha256:e635cf86354131f3b9e4872b8d5fc5170323f0224e439a443b2ff3d82d265a1c

Observation 19b5d2ad-b7c1-4368-8a6c-07d279cc7f99 · outbound

This paper cites an unresolved cited work.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Unresolved cited work

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T14:36:29.474796Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:36:29.474796Z digest=sha256:b168dfeb16a7bcf911733e05ba63fbd02df29d9f5812a71cb3cc2e82cfce3d3d

Observation c973e061-80aa-4b03-9b71-d641c42539c6 · outbound

This paper cites an unresolved cited work.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Unresolved cited work

Reference 41

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:36:31.138936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T14:36:29.532008Z digest=sha256:1de72efc346a6ecac549679ddcee27246dde66602a403f710d7628030f93fd06

Observation 9e461aa4-bd38-4e8b-a4bb-e5313aa4cb95 · outbound

This paper cites InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T14:36:29.583224Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:36:29.583224Z digest=sha256:0c5225ad47ec1f00abed70a7781db00e6fbea3760acd473bb8d09d25a299f346

Observation c9b73b3b-3978-44c9-8fe7-ea8da1860696 · outbound

This paper cites an unresolved cited work.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Unresolved cited work

Reference 43

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:36:31.450436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T14:36:29.509421Z digest=sha256:6044859c65b2f50e65a5f5b6589c2ed57d1efe7c1df540820e494c03e7fa3de7

Observation c5cf68d7-08d0-45c3-9875-4330fd4e649a · outbound

This paper cites an unresolved cited work.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Unresolved cited work

Reference 44

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:36:31.038003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T14:36:29.654739Z digest=sha256:24147580c3ba385934ade7293a16bffa3b4d668848698c5515808b3e3530fe0d

Observation 53ec807f-a90b-4c6d-a069-f22338da95a8 · outbound

This paper cites SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T14:36:29.673490Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:36:29.673490Z digest=sha256:cac34d147e289208e5e889e5d5ae88b8770c448e302df2198cd80adbeb5a5c7a

Observation 68df01c5-db36-4869-a1ba-cc133ae62b8f · outbound

This paper cites CLIP-ViP: Adapting Pre-trained Image-Text Model to Video-Language Representation Alignment.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning CLIP-ViP: Adapting Pre-trained Image-Text Model to Video-Language Representation Alignment

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T14:36:29.684895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:36:29.684895Z digest=sha256:6645822b59cf35559de022b71a644fd68d37f0efb056956ff48ccfd6bda97cb5

Observation 5835d953-a9c2-44f1-84de-92bb33d1d9e7 · outbound

This paper cites an unresolved cited work.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Unresolved cited work

Reference 47

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:36:31.110929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T14:36:29.607606Z digest=sha256:14ac1c39a92c41b346bef4e98b2dc65ac59da2bbddd78123af699eb4206a6164

Observation d7067fad-f68f-460a-a4fe-d845ebd89253 · outbound

This paper cites an unresolved cited work.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Unresolved cited work

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T14:36:29.709137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:36:29.709137Z digest=sha256:4e013104fe6bef70b6d9f55a7c0590a485569681f8dbb6a0be4cc49d27c0d705

Observation 31dcf7d2-1e4f-48d5-88ad-189e61538283 · outbound

This paper cites an unresolved cited work.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Unresolved cited work

Reference 49

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:36:30.976741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T14:36:29.726957Z digest=sha256:03f5dedee0457c4c454192fe6f535b3f6b2163d620a2b799c18f0beedd8d7f34

Observation 29c554df-1207-4a9c-80f4-2f8247fbcc36 · outbound

This paper cites an unresolved cited work.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Unresolved cited work

Reference 50

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:36:30.946168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T14:36:29.741989Z digest=sha256:83aedf902a1baac03706c0acccfdffcfec0b5b6dc4070e1a2e61ab3128c4b160

Observation 03d931d2-76c4-40a9-b6b1-26f88676df9e · outbound

This paper cites an unresolved cited work.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Unresolved cited work

Reference 51

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:36:30.906135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T14:36:29.748938Z digest=sha256:edd38d580a045a818326c3c215d5edd5db17277711a86f521a37985b45e1cdb7

Observation a70c7932-59e7-4ac5-a0af-b4d76ac58e99 · outbound

This paper cites Qwen3 Technical Report.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Qwen3 Technical Report

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T14:36:29.695562Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:36:29.695562Z digest=sha256:555e9ee8a033e563cab98ba1259f8c1335687f3e45e2afb26466e86a2c6ec0f2

Observation fcf9863a-fb78-41c6-92da-21c52c478f7c · outbound

This paper cites VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T14:36:29.767987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:36:29.767987Z digest=sha256:f5337387d882975bd15282c6e2ca776511b90f095997bf9c41d3af1aa4fc5b7c

Observation be0ef333-e251-4713-8cf3-b698da8e4b71 · outbound

This paper cites Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T14:36:29.714217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:36:29.714217Z digest=sha256:1d1e8ef37d47a0cbc5ef3abd5ee27a5f07d0c3c3f08848ac0fe4715e2dd765ef

Observation 918efbcc-c026-423b-8c60-ec56077fd46b · outbound

This paper cites an unresolved cited work.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Unresolved cited work

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T14:36:29.787485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:36:29.787485Z digest=sha256:1ca479d1ebd11df371791969215d31e52c9255397d21d03f172dcf2e3b64901e

Observation 5e30fd84-e35e-4b66-8987-d6987b31df39 · outbound

This paper cites an unresolved cited work.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Unresolved cited work

Reference 56

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:36:30.825970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T14:36:29.797201Z digest=sha256:3c4f63cb6487d8d67f7ee473250bb25293bc70b94c8c4627602e5e36c031d668

Observation 87f6d969-dcfa-47c6-a344-a8830c003976 · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-06T14:36:29.809595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:36:29.809595Z digest=sha256:76e1503b987a2df409bb6ab3874047fa86fdf6a3d359acfe5b54e777ba619d59

Observation 535d5528-c211-40a6-89a1-41072bdde2d6 · outbound

This paper cites Ferret: Refer and Ground Anything Anywhere at Any Granularity.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Ferret: Refer and Ground Anything Anywhere at Any Granularity

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-06T14:36:29.762895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:36:29.762895Z digest=sha256:45abf340821fa469347134089e17fcaa85139b5ed96019fa2b86ab0eacd1cf28

Observation 97411bd4-bb74-4ae5-b95c-028126faebd9 · outbound

This paper cites InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-06T14:36:29.862040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:36:29.862040Z digest=sha256:0f39da46aaf9dac75f45d954628d928e93f5997cc05a7f87febdd0440cd43470

Observation 03c16232-4046-467a-8847-08fb25ec7307 · outbound

This paper cites an unresolved cited work.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Unresolved cited work

Reference 60

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:36:30.864740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T14:36:29.780390Z digest=sha256:c53b90bc7d30200f79f4878c44481e8492edf4f9aa77476b63122ac5882d2fd3

Observation 3d196621-58ea-4326-ad9c-5b540ec82e8b · outbound

This paper cites OVC-Net: Object-Oriented Video Captioning with Temporal Graph and Detail Enhancement.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning OVC-Net: Object-Oriented Video Captioning with Temporal Graph and Detail Enhancement

Reference 64

Resolution
verified exact
local_arxiv, observed 2026-08-06T14:36:30.025967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T14:36:29.835315Z digest=sha256:f8197e0c2307203eef17d123c74201ee32bfa95924c39672c4e7d6b499ea6b4a

Observation 1c897235-9b2f-48d4-afe8-7c2a0937e9fc · outbound

This paper cites In Proceedings of the IEEE conference on computer vision and pattern recognition.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning In Proceedings of the IEEE conference on computer vision and pattern recognition

Reference 2018

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:36:31.241487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T14:36:29.515564Z digest=sha256:21a66cee1a71ec9ed394db62aac043953de44a7d5ef0da1b2b59155ed179c382

Observation 9e3a30e2-4a94-4ba2-abf5-c0628384966f · outbound

This paper cites In Proceedings of the IEEE/CVF international conference on computer vision.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning In Proceedings of the IEEE/CVF international conference on computer vision

Reference 2019

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:36:32.304827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T14:36:29.356331Z digest=sha256:c949e2fb7b3609c54c9a8f654793ac3b4b17d411585b70f8b36d7d2be82eb57a

Observation 4d657601-617a-4e66-80b6-9ff221e4d053 · outbound

This paper cites In 2020 IEEE International Conference on Multimedia and Expo (ICME).

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning In 2020 IEEE International Conference on Multimedia and Expo (ICME)

Reference 2020

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:36:31.068892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T14:36:29.625877Z digest=sha256:4598fe694bd435c1f108954e04aa20f0ea06509bb39147c011ce4a5472bed174

Observation a3898842-383a-46dc-bcdc-848142eec563 · outbound

This paper cites Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-06T14:36:28.852339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:36:28.852339Z digest=sha256:a0349701c78a077896a5db73e7d7b4c6ed8613499c5109da8ce3a54f201c5b83

Pith citing papers

Observation dc631b3f-aa4e-4c98-9e70-26fcb13e0df4 · inbound

RoadTones: Tone Controllable Text Generation from Road Event Videos cites this paper.

RoadTones: Tone Controllable Text Generation from Road Event Videos IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-21T04:39:35.270214Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T04:35:39.356919Z digest=sha256:e3c064676d229e25b2705d641f1acc5ae2f4c3246fe94d772efee57431cfa4d3