Pith. sign in

Paper Citation Record · LEDGER

Controllable Video Captioning with POS Sequence Guidance Based on Gated Fusion Network

As of 17 August 2026, this Paper Citation Record lists 64 of 64 outbound references and 0 inbound Pith citation observations for arXiv:1908.10072.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
1908.10072 v1

Coverage vector

measured 64 of 64 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-14T10:57:38.415005Z

measured 64 of 64 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

64 of 64 outbound references displayed

  • verified exact0
  • verified fuzzy53
  • unresolved11
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1873e6e7-db75-48f2-994e-6742a4b95208 · outbound

This paper cites Meteor: An automatic metric for mt evaluation with improved correlation with human judgments.

Controllable Video Captioning with POS Sequence Guidance Based on Gated Fusion Network Meteor: An automatic metric for mt evaluation with improved correlation with human judgments

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:57:39.651622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-14T10:57:38.041580Z digest=sha256:6f1d71ba72a447194cabac4b693cf2d01a7b9a65a0cb0801b5042d897d7c8b7e

Observation 021b16e0-d146-4917-b487-dd4c5234eea9 · outbound

This paper cites Quo vadis, action recognition? a new model and the kinetics dataset.

Controllable Video Captioning with POS Sequence Guidance Based on Gated Fusion Network Quo vadis, action recognition? a new model and the kinetics dataset

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:57:39.629639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-14T10:57:38.047580Z digest=sha256:0bf6b48ebb94ff4bf4507bb9e47a5dd07a6f27977eafbbab23b79270cd32d6f4

Observation c7d2996c-835a-41d4-ba78-3a57282b3503 · outbound

This paper cites Collecting highly parallel data for paraphrase evaluation.

Controllable Video Captioning with POS Sequence Guidance Based on Gated Fusion Network Collecting highly parallel data for paraphrase evaluation

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:57:39.612282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-14T10:57:38.053330Z digest=sha256:eed7a10772b88444770b8ad550d1f6c0999a38905d53fd7acf874b6b8a40fd3d

Observation 9df306c4-4e7e-4c5c-a4db-86ac34445603 · outbound

This paper cites Video captioning with guidance of multimodal latent topics.

Controllable Video Captioning with POS Sequence Guidance Based on Gated Fusion Network Video captioning with guidance of multimodal latent topics

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:57:39.594185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-14T10:57:38.059178Z digest=sha256:850de90812eb3d8b8e8bed86caf234969d58a09db0ad47b9fc538213c7b74820

Observation 41befd8a-49ea-4e0f-926d-ae5b944f8b9c · outbound

This paper cites Motion guided spatial attention for video captioning.

Controllable Video Captioning with POS Sequence Guidance Based on Gated Fusion Network Motion guided spatial attention for video captioning

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:57:39.573543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-14T10:57:38.064496Z digest=sha256:0e62138d85134191d407efe8d3a094787009dc4679c38ec4b55aacda77d9c703

Observation 7de028ad-7c2e-4fde-a3a9-8c9e009071cb · outbound

This paper cites Microsoft COCO Captions: Data Collection and Evaluation Server.

Controllable Video Captioning with POS Sequence Guidance Based on Gated Fusion Network Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-14T10:57:38.069975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T10:57:38.069975Z digest=sha256:c9c1beaa149bbe4043d6a9b63d255af6e8673d79aedc2616151cfc63d425145e

Observation 481f45ee-4ef8-412c-8145-ec313b212b6f · outbound

This paper cites Regularizing rnns for caption generation by reconstructing the past with the present.

Controllable Video Captioning with POS Sequence Guidance Based on Gated Fusion Network Regularizing rnns for caption generation by reconstructing the past with the present

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:57:39.554748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-14T10:57:38.076492Z digest=sha256:650b132e22d7e28efb877e45b44371877b7be704e7a7ce9f164eaa9151af919c

Observation d75d3be0-b701-4797-88fd-00040d78ea1e · outbound

This paper cites Less is more: Picking informative frames for video captioning.

Controllable Video Captioning with POS Sequence Guidance Based on Gated Fusion Network Less is more: Picking informative frames for video captioning

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:57:39.538263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-14T10:57:38.081380Z digest=sha256:5ef98dc2a63cd90cde4c1fd87b264a9cc0b2c82ac69c3060d5860ef2ebb553a8

Observation 34ce0a48-94d5-47ae-9fe5-34e52be0ad7d · outbound

This paper cites Embodied question answering.

Controllable Video Captioning with POS Sequence Guidance Based on Gated Fusion Network Embodied question answering

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:57:39.521481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-14T10:57:38.086671Z digest=sha256:c43aac63c830cf1892c573be7125ef8a7c1fb669fad7df5d7ad8f19b9535e4d9

Observation cab863b8-345a-49e2-95ac-b02654e6a053 · outbound

This paper cites Diverse and controllable image captioning with part-of-speech guidance.

Controllable Video Captioning with POS Sequence Guidance Based on Gated Fusion Network Diverse and controllable image captioning with part-of-speech guidance

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:57:39.504035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-14T10:57:38.091702Z digest=sha256:1b058e06bd412cf3658fd955baa60ed3efeecbfdeb1a6f016d445d50f4725fc4

Observation 2c82ff3d-679d-4f8d-ba5a-e0381c22deb6 · outbound

This paper cites Long-term recurrent convolutional networks for visual recognition and description.

Controllable Video Captioning with POS Sequence Guidance Based on Gated Fusion Network Long-term recurrent convolutional networks for visual recognition and description

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:57:39.487390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-14T10:57:38.099165Z digest=sha256:b8131bda99e7f35aa7a27e90de1cff00fe24bd5d3a27899936f4685f8f27f9b2

Observation f8fd452f-6470-4535-bc1a-c6e806eb8b79 · outbound

This paper cites Unsupervised image captioning.

Controllable Video Captioning with POS Sequence Guidance Based on Gated Fusion Network Unsupervised image captioning

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:57:39.468755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-14T10:57:38.104751Z digest=sha256:6665e3be72f5f5d3de3c787ea173c408f6d2ceb833c570b26a82313b58803b63

Observation 5091e6f7-4a65-40e0-ab93-8ad3849c18b5 · outbound

This paper cites Multimodal compact bilinear pooling for visual question answering and visual grounding.

Controllable Video Captioning with POS Sequence Guidance Based on Gated Fusion Network Multimodal compact bilinear pooling for visual question answering and visual grounding

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:57:39.450592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-14T10:57:38.109755Z digest=sha256:c43d596afe1b00e5f336a601efbcf7f60a2311a6f2491beef18e8112bb41f382

Observation 4df725c0-297e-43d6-b786-a5bca031cd73 · outbound

This paper cites Semantic compositional networks for visual captioning.

Controllable Video Captioning with POS Sequence Guidance Based on Gated Fusion Network Semantic compositional networks for visual captioning

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:57:39.433743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-14T10:57:38.115264Z digest=sha256:8afcbd836a118992489a511b40abb3188ba17dcdffc8e825258093a5192ae596

Observation 5f1b3ffe-4ac0-4288-9a15-f6affc13d424 · outbound

This paper cites Youtube2text: Recognizing and describing arbitrary activities using semantic hierarchies and zero-shot recognition.

Controllable Video Captioning with POS Sequence Guidance Based on Gated Fusion Network Youtube2text: Recognizing and describing arbitrary activities using semantic hierarchies and zero-shot recognition

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:57:39.414528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-14T10:57:38.121213Z digest=sha256:dd94d5dfa7970bcf9e1869dd927d2999385ec18d8760783da81970298a19593d

Observation 3a93c80b-12b3-45ec-86a7-4ca44ba535a6 · outbound

This paper cites Image caption generation with part of speech guidance.

Controllable Video Captioning with POS Sequence Guidance Based on Gated Fusion Network Image caption generation with part of speech guidance

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:57:39.392423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-14T10:57:38.126470Z digest=sha256:8538b702232b76caef69fdd55a33d8378d7bcd532236b91f269b7af302d6dc26

Observation d137cd4b-5e2f-46a4-b370-e29aa33a6ecd · outbound

This paper cites Recurrent fusion network for image captioning.

Controllable Video Captioning with POS Sequence Guidance Based on Gated Fusion Network Recurrent fusion network for image captioning

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:57:39.373534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-14T10:57:38.131519Z digest=sha256:e0ba8e178c626c8bc9b62a377a9fd720cef32bf106852972765b1068712f6c49

Observation 37133a6b-b580-46ff-9cf8-4f59c9d78d39 · outbound

This paper cites Describing videos using multi-modal fusion.

Controllable Video Captioning with POS Sequence Guidance Based on Gated Fusion Network Describing videos using multi-modal fusion

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:57:39.354958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-14T10:57:38.137521Z digest=sha256:e50bca4179a0988c971c1d9b26c32d1743d856e275ee05935a1e596756cd42fb

Observation b11ab0e7-a50b-4e99-af94-a08eca01a820 · outbound

This paper cites The Kinetics Human Action Video Dataset.

Controllable Video Captioning with POS Sequence Guidance Based on Gated Fusion Network The Kinetics Human Action Video Dataset

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-14T10:57:38.143226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T10:57:38.143226Z digest=sha256:5066dc376d0c2e1da08cde0277acb4ae404ff63232d4269ca025f901d51f4c33

Observation c7fcec0e-8678-4bc7-ae8b-d7fee8afe461 · outbound

This paper cites Hadamard product for low-rank bilinear pooling.

Controllable Video Captioning with POS Sequence Guidance Based on Gated Fusion Network Hadamard product for low-rank bilinear pooling

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:57:39.337697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-14T10:57:38.151471Z digest=sha256:d7cef312c4d6e7a2820209b88d4eb9f0aae12ed2eb2541b13eddd12cdea1487e

Observation fd5448d2-06ec-4839-94be-8a646803e163 · outbound

This paper cites Natural language description of human activities from video images based on concept hierarchy of actions.

Controllable Video Captioning with POS Sequence Guidance Based on Gated Fusion Network Natural language description of human activities from video images based on concept hierarchy of actions

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:57:39.319316Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-14T10:57:38.157564Z digest=sha256:538be6130a22054f4619a9064d2f4de6f382cafe7ae17ed09927ec16e2aec834

Observation 5baa7aad-7032-48e8-aa80-26f5f6fc6183 · outbound

This paper cites Rouge: A package for automatic evaluation of summaries.

Controllable Video Captioning with POS Sequence Guidance Based on Gated Fusion Network Rouge: A package for automatic evaluation of summaries

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-14T10:57:38.168749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T10:57:38.168749Z digest=sha256:63910ac4e114a7296ac99d416b8a6d81862104c39f0f9e65d7f53af07d289efd

Observation d6ffe39b-b72a-4915-8561-e5da7eb56eaf · outbound

This paper cites Sibnet: Sibling convolutional encoder for video captioning.

Controllable Video Captioning with POS Sequence Guidance Based on Gated Fusion Network Sibnet: Sibling convolutional encoder for video captioning

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:57:39.279723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-14T10:57:38.176170Z digest=sha256:38c0e60069ab0af18bb9d9a6bae8087e84ae94fec448cc02a342cebbe0e02b60

Observation 1bc14e2a-7463-42a6-a223-3813c20f5cce · outbound

This paper cites Matching image and sentence with multi-faceted representations.

Controllable Video Captioning with POS Sequence Guidance Based on Gated Fusion Network Matching image and sentence with multi-faceted representations

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:57:39.260563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-14T10:57:38.181550Z digest=sha256:ba4000a1f9f2b20a421f42f5f4af9a1bdcb506b12bf7117df18a9a14b1f83cf9

Observation 700d4d2c-c917-4829-bd71-ea8df1bb4df0 · outbound

This paper cites Learning to answer questions from image using convolutional neural network.

Controllable Video Captioning with POS Sequence Guidance Based on Gated Fusion Network Learning to answer questions from image using convolutional neural network

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:57:39.244241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-14T10:57:38.186885Z digest=sha256:b2fb2d3811a85070e336c6d00e8c3f8107cf36efd9389273ea57b02ccbdbc59e

Observation 9a572c67-6bea-4cd4-be14-2846aa3b6a65 · outbound

This paper cites Multimodal convolutional neural networks for matching image and sentence.

Controllable Video Captioning with POS Sequence Guidance Based on Gated Fusion Network Multimodal convolutional neural networks for matching image and sentence

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:57:39.227761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-14T10:57:38.192462Z digest=sha256:83a620fb4e9a47150fcca2650bf546cf0b4921173c91a7781399eb7c0e4a0e7d

Observation c9dae1cb-c071-497e-8df2-ade0bfbab609 · outbound

This paper cites Jointly modeling embedding and translation to bridge video and language.

Controllable Video Captioning with POS Sequence Guidance Based on Gated Fusion Network Jointly modeling embedding and translation to bridge video and language

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:57:39.211142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-14T10:57:38.198856Z digest=sha256:6ce7e21744005556cefe7b2226a7df40be5a4fc7124eab9fe6c07a7b1b8a6207

Observation c5c7377c-0eea-44d0-9384-882e23422678 · outbound

This paper cites Video captioning with transferred semantic attributes.

Controllable Video Captioning with POS Sequence Guidance Based on Gated Fusion Network Video captioning with transferred semantic attributes

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:57:39.192777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-14T10:57:38.204428Z digest=sha256:461342a1828c4cd84a3ea817114422dc60b4d67654f5cc105ccd653670d430f9

Observation 17f6b06a-f575-46f8-9748-75a7407ac00d · outbound

This paper cites Bleu: a method for automatic evaluation of machine translation.

Controllable Video Captioning with POS Sequence Guidance Based on Gated Fusion Network Bleu: a method for automatic evaluation of machine translation

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:57:39.176958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-14T10:57:38.209681Z digest=sha256:0c2b044275372c31f8133b225b982bbced565b96176f4dceaec6307df04bd165

Observation d16035c1-df84-4eb6-9130-bbee5849f4fc · outbound

This paper cites Tv-l1 optical flow estimation.

Controllable Video Captioning with POS Sequence Guidance Based on Gated Fusion Network Tv-l1 optical flow estimation

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:57:39.159957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-14T10:57:38.215643Z digest=sha256:771e9bf960da64b9e3a58fe2df9ecb5cbdb3ed402431e3171f10ce98a8c68aeb

Observation 1f5c2d93-c0ba-4599-8f61-dec9b559090e · outbound

This paper cites Multimodal video description.

Controllable Video Captioning with POS Sequence Guidance Based on Gated Fusion Network Multimodal video description

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:57:39.143375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-14T10:57:38.220829Z digest=sha256:b19045389f75da039a699d438b59bf50d1747510de77ca90f4d70e6abc424b13

Observation 125438e3-74ba-4c9c-a42c-ed07a30c43a5 · outbound

This paper cites Self-critical sequence training for image captioning.

Controllable Video Captioning with POS Sequence Guidance Based on Gated Fusion Network Self-critical sequence training for image captioning

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:57:39.126286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-14T10:57:38.226951Z digest=sha256:2cdf6043dcbcc07447a2ef16aab4afb7aca68cc753884fce6983abf33b390b80

Observation 0a6eba4c-8f05-4e06-99ed-98f2b8e01e39 · outbound

This paper cites Coherent multi-sentence video description with variable level of detail.

Controllable Video Captioning with POS Sequence Guidance Based on Gated Fusion Network Coherent multi-sentence video description with variable level of detail

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:57:39.108828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-14T10:57:38.232533Z digest=sha256:ea6cc43677bd1160268ea3044ca15eccba1eedb7f7d2e92b5b7e9188e1a83758

Observation 25f0ed17-dd4c-440b-b6ac-b34f75bacfd9 · outbound

This paper cites Translating video content to natural language descriptions.

Controllable Video Captioning with POS Sequence Guidance Based on Gated Fusion Network Translating video content to natural language descriptions

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:57:39.093009Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-14T10:57:38.238054Z digest=sha256:66295436afd065d6b31b5382df5d046426a84b8710f60a73700877856f9c7e92

Observation 1bb049a4-948f-4de3-8b78-667998310f40 · outbound

This paper cites Imagenet large scale visual recognition challenge.

Controllable Video Captioning with POS Sequence Guidance Based on Gated Fusion Network Imagenet large scale visual recognition challenge

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:57:39.075893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-14T10:57:38.246095Z digest=sha256:817630d461342f796930eebe08889e832fa87fbbd05652107ee941df2b2cdb34

Observation 5d73ab8b-b969-4a39-a3b5-52386a26d2d1 · outbound

This paper cites Frame-and segment-level features and candidate pool evaluation for video caption generation.

Controllable Video Captioning with POS Sequence Guidance Based on Gated Fusion Network Frame-and segment-level features and candidate pool evaluation for video caption generation

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:57:39.058047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-14T10:57:38.252404Z digest=sha256:7c630963ebba844093c06494df53b54ef40304d3eb7611c278b3b1dcc107af06

Observation c7427a89-25d2-4d56-ac26-5c1b8981e5b1 · outbound

This paper cites Quantization-based hashing: a general framework for scalable image and video retrieval.

Controllable Video Captioning with POS Sequence Guidance Based on Gated Fusion Network Quantization-based hashing: a general framework for scalable image and video retrieval

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:57:39.039809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-14T10:57:38.258742Z digest=sha256:b0226402023ad822ae7326d507bf9a6f5cc6ab1a315a800d8107e06147e26b38

Observation 328ed9a7-2499-4e16-8773-4450a66adadf · outbound

This paper cites Inception-v4, inception-resnet and the impact of residual connections on learning.

Controllable Video Captioning with POS Sequence Guidance Based on Gated Fusion Network Inception-v4, inception-resnet and the impact of residual connections on learning

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:57:39.022436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-14T10:57:38.265410Z digest=sha256:d332b2c99431dd2716340bc02248c78e8f9fc487aed0c86794090f52f3f54220

Observation a8ef05c7-5dff-4789-9594-33bf05158b2c · outbound

This paper cites Feature-rich part-of-speech tagging with a cyclic dependency network.

Controllable Video Captioning with POS Sequence Guidance Based on Gated Fusion Network Feature-rich part-of-speech tagging with a cyclic dependency network

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:57:39.003880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-14T10:57:38.272559Z digest=sha256:90129610214e9fd4aa6f144feafad35407355541cbc7a1a95ec95042b3e6bc9c

Observation f73729ed-960a-4aea-9987-e82694235fea · outbound

This paper cites Learning spatiotemporal features with 3d convolutional networks.

Controllable Video Captioning with POS Sequence Guidance Based on Gated Fusion Network Learning spatiotemporal features with 3d convolutional networks

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:57:38.981064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-14T10:57:38.279403Z digest=sha256:6a509f1443e73cabfa2ae790746b95d485c731c50ed22f93f5d9098350fa601e

Observation 9e318745-2eb7-46a5-bc12-d5c4a9cc3744 · outbound

This paper cites Cider: Consensus-based image description evaluation.

Controllable Video Captioning with POS Sequence Guidance Based on Gated Fusion Network Cider: Consensus-based image description evaluation

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:57:38.957250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-14T10:57:38.284378Z digest=sha256:bb7ab10d2e0b3e0ce4535524c14d847bb55054c1cd442e5c47d8f83920858f80

Observation c4fb9dbb-b469-4a6e-a081-6a7517c5fc11 · outbound

This paper cites Sequence to sequence-video to text.

Controllable Video Captioning with POS Sequence Guidance Based on Gated Fusion Network Sequence to sequence-video to text

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:57:38.940058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-14T10:57:38.290185Z digest=sha256:db52f03f555326d6f0506dbc553eb1bec7e5d07754e7523fbf48d4b3780840b0

Observation e5e326db-5329-435d-8d2e-2e5f3a6da194 · outbound

This paper cites Translating videos to natural language using deep recurrent neural networks.

Controllable Video Captioning with POS Sequence Guidance Based on Gated Fusion Network Translating videos to natural language using deep recurrent neural networks

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:57:38.921555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-14T10:57:38.297238Z digest=sha256:ad9d70f7388aaf057818afb95a27941a470fdfd4a781bfd27f0be28a1fd7bf8b

Observation 408adeae-2df5-4cde-994f-197dd7e1f8e2 · outbound

This paper cites Hierarchical photo-scene encoder for album storytelling.

Controllable Video Captioning with POS Sequence Guidance Based on Gated Fusion Network Hierarchical photo-scene encoder for album storytelling

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:57:38.898224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-14T10:57:38.303765Z digest=sha256:f51def1a21b7fe306da17579ee2755e0470b32021e55b43500e144c6fc10239a

Observation c28207be-6357-408c-9aee-dd507a4f322e · outbound

This paper cites Reconstruction network for video captioning.

Controllable Video Captioning with POS Sequence Guidance Based on Gated Fusion Network Reconstruction network for video captioning

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:57:38.879932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-14T10:57:38.309023Z digest=sha256:8fc4ec76a7301c79ea33284c4c72feb35843a6935b1746d2ff1ea7b1ee66469c

Observation 0da1353e-262b-4929-8318-89f3cf4832e4 · outbound

This paper cites Bidirectional attentive fusion with context gating for dense video captioning.

Controllable Video Captioning with POS Sequence Guidance Based on Gated Fusion Network Bidirectional attentive fusion with context gating for dense video captioning

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:57:38.861169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-14T10:57:38.314056Z digest=sha256:14e9a21dcc74a83b1a66e51d6d62806b4e552acf9aaf7ec31848b7fb9c79f2ec

Observation fb9e6b81-abac-4d4c-b220-38e3d1fa02fa · outbound

This paper cites M3: Multimodal memory modelling for video captioning.

Controllable Video Captioning with POS Sequence Guidance Based on Gated Fusion Network M3: Multimodal memory modelling for video captioning

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:57:38.843411Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-14T10:57:38.319099Z digest=sha256:b4fff6e0e471796aace841e05a842964a25dfa0828691207e12e1b930dcfe146

Observation f54d4121-39f4-4b6a-8d35-7ff8fae36104 · outbound

This paper cites A survey on learning to hash.

Controllable Video Captioning with POS Sequence Guidance Based on Gated Fusion Network A survey on learning to hash

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:57:38.825218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-14T10:57:38.324097Z digest=sha256:b80ab0b0a93783cee1f9d8b9c4574179aa4e756f71751083302c23e11ba24ad2

Observation 8a66c135-e4c8-4d55-b55b-1cd98878b26f · outbound

This paper cites Interpretable video captioning via trajectory structured localization.

Controllable Video Captioning with POS Sequence Guidance Based on Gated Fusion Network Interpretable video captioning via trajectory structured localization

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:57:38.807467Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-14T10:57:38.329406Z digest=sha256:691d0b15d93b60042ee3201e5f56e28764cef2c343db9c9409ab22ba4bb03c98

Observation 28b260c5-3bdc-4cb1-abaa-56167eece4f1 · outbound

This paper cites Msr-vtt: A large video description dataset for bridging video and language.

Controllable Video Captioning with POS Sequence Guidance Based on Gated Fusion Network Msr-vtt: A large video description dataset for bridging video and language

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:57:38.788354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-14T10:57:38.334432Z digest=sha256:9d9f1866c013a0b4e391754fc2824b0770dfd907ae08c58355dc92459835b07d

Observation 50162c04-e9e6-4aa3-85f5-32c46c1ffb3d · outbound

This paper cites Learning multimodal attention lstm networks for video captioning.

Controllable Video Captioning with POS Sequence Guidance Based on Gated Fusion Network Learning multimodal attention lstm networks for video captioning

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:57:38.765691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-14T10:57:38.339586Z digest=sha256:f43e4ae019d6f5d1be73f5ae8afc0c0fa3bce0ba4c034f6ff5c97153b3bd81b8

Observation 3f0aaf19-3837-458e-9e34-87f77d321aa7 · outbound

This paper cites Jointly modeling deep video and compositional text to bridge vision and language in a unified framework.

Controllable Video Captioning with POS Sequence Guidance Based on Gated Fusion Network Jointly modeling deep video and compositional text to bridge vision and language in a unified framework

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:57:38.746357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-14T10:57:38.345161Z digest=sha256:aeaf2d93763173a0ed3756f378a2e7d298eefb5ed78bc5ea53486ca106cbc17f

Observation 658ae573-f81c-4559-bda4-e7e81f1d461b · outbound

This paper cites Describing videos by exploiting temporal structure.

Controllable Video Captioning with POS Sequence Guidance Based on Gated Fusion Network Describing videos by exploiting temporal structure

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:57:38.724518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-14T10:57:38.354088Z digest=sha256:135091e285ee71e58f2cc8de988935bdf5c7f00ab819d07c786061eae75a495c

Observation a010624a-10c4-45f7-b188-3f0f6482ef5b · outbound

This paper cites Video paragraph captioning using hierarchical recurrent neural networks.

Controllable Video Captioning with POS Sequence Guidance Based on Gated Fusion Network Video paragraph captioning using hierarchical recurrent neural networks

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:57:38.703201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-14T10:57:38.360033Z digest=sha256:d49ff678afdef5336c28e6d0b59ef2bc2e6fd348892a9659dc2ab10a15f37aa2

Observation efd89af8-df75-4e7f-b627-91de1c30eefc · outbound

This paper cites ADADELTA: An Adaptive Learning Rate Method.

Controllable Video Captioning with POS Sequence Guidance Based on Gated Fusion Network ADADELTA: An Adaptive Learning Rate Method

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-14T10:57:38.365230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T10:57:38.365230Z digest=sha256:4644fc89000f881073c062fddb566c6f66d8e6501ffaaafbcd1157a42f02b555

Observation 4296d482-517e-4e48-964d-212286737a98 · outbound

This paper cites Reconstruct and represent video contents for captioning via reinforcement learning.

Controllable Video Captioning with POS Sequence Guidance Based on Gated Fusion Network Reconstruct and represent video contents for captioning via reinforcement learning

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:57:38.680241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-14T10:57:38.371492Z digest=sha256:d3dd228a96d9fc90c2f807121a9d04c30f2224885b742f223daad567b7b89ba0

Observation cab27f6e-071f-41d5-9e55-f7c21fd2b2b2 · outbound

This paper cites an unresolved cited work.

Controllable Video Captioning with POS Sequence Guidance Based on Gated Fusion Network Unresolved cited work

Reference 57

Resolution
unresolved
raw_fallback, observed 2026-08-14T10:57:38.656895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-14T10:57:38.376450Z digest=sha256:47455f72fd9f68019e604bea244a1c6587967fea649c5a76d7ab7a3f3a6f0061

Observation 9575ebe7-20b6-4b41-90cd-6d53b600021d · outbound

This paper cites Reinforced Video Captioning with Entailment Rewards.

Controllable Video Captioning with POS Sequence Guidance Based on Gated Fusion Network Reinforced Video Captioning with Entailment Rewards

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-14T10:57:38.381673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T10:57:38.381673Z digest=sha256:695ca3ce0bfb95219e2bbea35b3799db89b2149d2b9760aa5630f39fb5903d67

Observation 18fdd141-200f-401e-9a10-76be6874bb12 · outbound

This paper cites an unresolved cited work.

Controllable Video Captioning with POS Sequence Guidance Based on Gated Fusion Network Unresolved cited work

Reference 59

Resolution
unresolved
raw_fallback, observed 2026-08-14T10:57:38.636429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-14T10:57:38.387488Z digest=sha256:3a42d125679205b6f6204fcc80d056bf21c71bc39318421239c99758b4015413

Observation 71e4be24-20f1-4bf1-aaba-8833c9e222e4 · outbound

This paper cites an unresolved cited work.

Controllable Video Captioning with POS Sequence Guidance Based on Gated Fusion Network Unresolved cited work

Reference 60

Resolution
unresolved
raw_fallback, observed 2026-08-14T10:57:38.618714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-14T10:57:38.393036Z digest=sha256:b09f607769e852cae492a0d2a0c43e99093dd903d364704dee26eb0a17bc0348

Observation 5d62de1a-c435-4887-85d2-63b8ad041f82 · outbound

This paper cites an unresolved cited work.

Controllable Video Captioning with POS Sequence Guidance Based on Gated Fusion Network Unresolved cited work

Reference 61

Resolution
unresolved
raw_fallback, observed 2026-08-14T10:57:38.600523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-14T10:57:38.398564Z digest=sha256:55c549bd352a05bdd3ded2279141fd718d741f37b817db30382b5428415d61e6

Observation c8db4f33-38ba-4b64-afa1-e2ab34ba190f · outbound

This paper cites Vedantam, C.

Controllable Video Captioning with POS Sequence Guidance Based on Gated Fusion Network Vedantam, C

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:57:38.582150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-14T10:57:38.404074Z digest=sha256:64ba4617cfaa7d47ca6fe0e6dabb255d86b932332902e410708b3936d6e70def

Observation 5a5ab054-33a5-47b3-a10b-26015aab2806 · outbound

This paper cites an unresolved cited work.

Controllable Video Captioning with POS Sequence Guidance Based on Gated Fusion Network Unresolved cited work

Reference 63

Resolution
unresolved
raw_fallback, observed 2026-08-14T10:57:38.561785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-14T10:57:38.409600Z digest=sha256:b8d5c501244bc0576ece4ed813245da43a855aac98faa975a1fc9f2c1d9d913f

Observation 37f52123-9752-4d5d-9975-2604d131c8bb · outbound

This paper cites Reinforcement Learning Neural Turing Machines - Revised.

Controllable Video Captioning with POS Sequence Guidance Based on Gated Fusion Network Reinforcement Learning Neural Turing Machines - Revised

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-14T10:57:38.415005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T10:57:38.415005Z digest=sha256:c78ba0e90e4b85e091bda18af47be0462b97a08b319ea65234ac315df97aaa56

Pith citing papers

No inbound Pith citation observations are available.