Pith. sign in

Paper Citation Record · LEDGER

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning

As of 15 August 2026, this Paper Citation Record lists 77 of 77 outbound references and 0 inbound Pith citation observations for arXiv:2411.14688.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.14688 v1

Coverage vector

measured 77 of 77 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T15:06:05.171159Z

measured 77 of 77 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

77 of 77 outbound references displayed

  • verified exact4
  • verified fuzzy47
  • unresolved25
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9fc1bf43-b827-4330-835c-6a2b45ab28c6 · outbound

This paper cites Flamingo: a visual language model for few-shot learning,.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Flamingo: a visual language model for few-shot learning,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T15:06:04.828160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:06:04.828160Z digest=sha256:cfc408deeaf11dd44114cda9ad9981df4fce8bb41b43aa4b76b572569621edbe

Observation 58eda270-95ed-4593-814f-882abfb3470e · outbound

This paper cites an unresolved cited work.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Unresolved cited work

Reference 2

Resolution
malformed identifier
raw_fallback, observed 2026-08-12T15:06:06.426126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T15:06:04.833355Z digest=sha256:150a492426cf8c7bceac752c3aa44800d2f4fe33e37e35bf5693b11985b92ef2

Observation c5f0066a-1638-40ab-93fa-266c263e0cce · outbound

This paper cites Meteor: An automatic metric for mt evaluation with improved correlation with hu- man judgments.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Meteor: An automatic metric for mt evaluation with improved correlation with hu- man judgments

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T15:06:04.837965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:06:04.837965Z digest=sha256:54483ed62fcf5ddde578836bf8394eec2cb490afab1623333020cfa2098126cc

Observation 1281e890-ec4b-4755-ae78-099640dadade · outbound

This paper cites BEiT: BERT Pre-Training of Image Transformers.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning BEiT: BERT Pre-Training of Image Transformers

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T15:06:04.842945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:06:04.842945Z digest=sha256:b491c154f02d4e5716b749a47a33857d7d371df7f275f15b0b951468efecc606

Observation 3f418b2a-b1fa-4535-bbab-388e3ab1876d · outbound

This paper cites Recur- rent memory transformer.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Recur- rent memory transformer

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:06:06.402403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T15:06:04.847643Z digest=sha256:5f2bfda68b0f88010ce7575125cde4925d5c566c229c050e6f1f2c9efb94c1d0

Observation 4b1e360d-70a1-4761-97a5-f603fb9458dc · outbound

This paper cites Quo vadis, action recognition? a new model and the kinetics dataset.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Quo vadis, action recognition? a new model and the kinetics dataset

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T15:06:04.852261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:06:04.852261Z digest=sha256:38c7f931dbeafd6000846eaab79d0088f0f3ce444cb060509a261095900f44d5

Observation 0933075a-1a6f-4b96-9f7e-6a0273406282 · outbound

This paper cites PaLI-X: On Scaling up a Multilingual Vision and Language Model.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning PaLI-X: On Scaling up a Multilingual Vision and Language Model

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T15:06:04.856994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:06:04.856994Z digest=sha256:b2e953df862bbdd59fc2201923f3a27fe63d20f78bc0e1cf8ae0437bc0d723ad

Observation 61cdd851-19ad-4708-b862-1df2747700ec · outbound

This paper cites PaLI: A jointly-scaled multilingual language- image model.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning PaLI: A jointly-scaled multilingual language- image model

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:06:06.378746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T15:06:04.861858Z digest=sha256:e6957706222c205bb7ec920feb064de07b24e327bb060c7aa193ebdbfc76282c

Observation 322233ca-9b62-4e1c-8271-0d32a8f6c89e · outbound

This paper cites VideoOFA: Two-Stage Pre-Training for Video-to-Text Generation.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning VideoOFA: Two-Stage Pre-Training for Video-to-Text Generation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T15:06:04.866401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:06:04.866401Z digest=sha256:3b84e34fb9b60397ad27f40709854ce10dd0b0c52f411c4069076cb70724665d

Observation 476c7f7b-ebdf-4ebc-b6aa-c61fb20c7ffa · outbound

This paper cites Uniter: Universal image-text representation learning.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Uniter: Universal image-text representation learning

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:06:06.364399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T15:06:04.871346Z digest=sha256:59ccc62e6692402bdb9c6419011b30371692a7663167f0df62ebb2475c4d7062

Observation db2c8653-8ef1-4f0a-8d31-e135ea8b7ea5 · outbound

This paper cites TALLFormer: Temporal Action Localization with a Long-memory Transformer.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning TALLFormer: Temporal Action Localization with a Long-memory Transformer

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-08-12T15:06:05.606074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T15:06:04.875901Z digest=sha256:3f7b2866ac733039c11e4e9e6fbb2068a8747349504134f6921b24f447b47740

Observation 62d8344d-51fd-4c6b-bae4-521dc7e62b3d · outbound

This paper cites Monotonic chunkwise attention.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Monotonic chunkwise attention

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:06:06.349731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T15:06:04.881088Z digest=sha256:501b5c5d3bc3f38ad4787ce189a8bca7c00e756396a33e7771250a1fffb4a81f

Observation 2f2ed543-2ca6-41ab-bcd2-41e67c5c429a · outbound

This paper cites Vision Transformers Need Registers.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Vision Transformers Need Registers

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T15:06:04.885366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:06:04.885366Z digest=sha256:f0e008135ac127d4674075184242e4807bbf447fd18163df2c65f29cb9ccee01

Observation 5b6b29fb-8f0a-4264-899a-69afab6caf9c · outbound

This paper cites An empirical study of training end-to-end vision-and-language transformers.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning An empirical study of training end-to-end vision-and-language transformers

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:06:06.335076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T15:06:04.890245Z digest=sha256:8b43da7cc603afcd4f2b467b485116e85d762810d47b2284c4f6e3f72b3eddae

Observation 64219baf-12c1-41bb-b565-5477d750100b · outbound

This paper cites Violet: End-to-end video-language transformers with masked visual-token mod- eling.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Violet: End-to-end video-language transformers with masked visual-token mod- eling

Reference 15

Resolution
verified exact
raw_fallback, observed 2026-08-12T15:06:05.568804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T15:06:04.894546Z digest=sha256:70cffeb7cf0cb165444aef4e1143944d152b3c86b8cf2c0bc36d0a3ccd1f904c

Observation 8dd6f2c7-edaa-4208-9707-51c3bd6387fb · outbound

This paper cites Soda: Story oriented dense video captioning evaluation framework.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Soda: Story oriented dense video captioning evaluation framework

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T15:06:04.898868Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:06:04.898868Z digest=sha256:f272ddc17446b2727d397b0d4de37a28e82ac95f20daebda4a6c80443bb0e9d2

Observation bf836ebb-e58f-4baf-9c41-2d409e04da11 · outbound

This paper cites Mist: Multi-modal iterative spatial- temporal transformer for long-form video question answer- ing.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Mist: Multi-modal iterative spatial- temporal transformer for long-form video question answer- ing

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:06:06.310920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T15:06:04.903286Z digest=sha256:b1d9e102883ae57f5ba46ea3c0a6451d69abde6ddfda7059e227b232900a49ce

Observation 779e2d54-5c99-4f82-aec2-59951c501850 · outbound

This paper cites The” something something” video database for learning and evaluating visual common sense.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning The” something something” video database for learning and evaluating visual common sense

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:06:06.296440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T15:06:04.907497Z digest=sha256:79888ce7bb965222964df2ef1b10621ea250fb94656ba751ace06d38ab135a13

Observation cb1b91ff-dee4-45ba-bdee-9f5c0a287456 · outbound

This paper cites Videollm: Modeling video sequence with large language models.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Videollm: Modeling video sequence with large language models

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:06:06.282074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T15:06:04.911726Z digest=sha256:1b01cb634dd1e1f2eebae6da9199060f92d891802d3c2370d1d348763de86264

Observation eb4d5fbf-0614-410f-b635-255f322addb1 · outbound

This paper cites Multimodal pretraining for dense video cap- tioning.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Multimodal pretraining for dense video cap- tioning

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:06:06.267858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T15:06:04.916036Z digest=sha256:429c931d92835dac21531fb439986bef3fad198217b8b8844f213cd9d548def1

Observation e534e57a-ec11-4924-91e7-549b3c17c99b · outbound

This paper cites A better use of audio-visual cues: Dense video captioning with bi-modal transformer.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning A better use of audio-visual cues: Dense video captioning with bi-modal transformer

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:06:06.253553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T15:06:04.920357Z digest=sha256:194e5af8d4be88f213dc6ed4e09568f82b4876ee71b06941d1f3dbef6b66c992

Observation 28f5fd1d-56f8-48f4-a619-f5a5ca7e567b · outbound

This paper cites Long movie clip classification with state-space video models.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Long movie clip classification with state-space video models

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:06:06.238572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T15:06:04.924749Z digest=sha256:a4de3ee125e3d618b885c3876567b9b301d042698aa9540e4ccd390aa5a86881

Observation 2ebbb894-1c9c-47ea-a25e-cc09b23d77d6 · outbound

This paper cites Perceiver: General perception with iterative attention, 2021.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Perceiver: General perception with iterative attention, 2021

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:06:06.223759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T15:06:04.929930Z digest=sha256:3c20e23721557db1f7487c14e22e3172af05eccfa514089c8e251c30d57e2e5a

Observation 866920ec-e48c-4654-a1d0-197ef9477bd4 · outbound

This paper cites Scaling up visual and vision-language representation learning with noisy text supervision.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Scaling up visual and vision-language representation learning with noisy text supervision

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:06:06.207617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T15:06:04.934248Z digest=sha256:110aebc71c7186cdf4eda9816f0be3f6338169509e61f03da5b86469ff6b1142

Observation 634c3a2a-8ba3-4870-b3e7-b5d1ada9c606 · outbound

This paper cites The Kinetics Human Action Video Dataset.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning The Kinetics Human Action Video Dataset

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T15:06:04.938660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:06:04.938660Z digest=sha256:4d65b2b124360bb81233eb9f9697b5838c0af9d5b786678556f4a06782206e87

Observation 7b2d3f33-9c37-433f-b9a0-d5471995577a · outbound

This paper cites Dense-captioning events in videos.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Dense-captioning events in videos

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:06:06.192583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T15:06:04.943485Z digest=sha256:39fd78e27a833181bf32b88838e080bcb359c815a9f7e26c535e532d99b528f8

Observation ab86b914-033e-4fe5-86ec-c876dbe04c34 · outbound

This paper cites MaMMUT: A simple architecture for joint learning for mul- timodal tasks.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning MaMMUT: A simple architecture for joint learning for mul- timodal tasks

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:06:06.177370Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T15:06:04.948145Z digest=sha256:2b95769735551baa113d718f2b12cb18d2942c249c52d4661e3b8ae553f3bddc

Observation 93a06bfb-34e4-4d4f-81f5-7c577724b57e · outbound

This paper cites Selvaraju, Akhilesh D.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Selvaraju, Akhilesh D

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:06:06.162380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T15:06:04.952500Z digest=sha256:2eb750a6f6b6c466fa6a634bd13560f9dcb21e47b31bcb5bc55f6ded352fc205

Observation 2073b075-6195-4090-a649-687ad2ea368c · outbound

This paper cites BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T15:06:04.956799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:06:04.956799Z digest=sha256:00b3ea4fbb338d7c60ddebe17ab0f2350b0e4f29c0c7cb9c13796f72fd9fdb12

Observation 16497643-3609-4795-953f-103946ee1173 · outbound

This paper cites Unmasked teacher: Towards training-efficient video foundation models.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Unmasked teacher: Towards training-efficient video foundation models

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:06:06.146858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T15:06:04.961415Z digest=sha256:1f7815688663fc21336522412302dbff27a11b18d062a4731b67619d1eb39eb3

Observation ccea7bdf-7641-42e0-b7ec-5256e04f62a9 · outbound

This paper cites Oscar: Object-semantics aligned pre-training for vision-language tasks.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Oscar: Object-semantics aligned pre-training for vision-language tasks

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T15:06:04.965854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:06:04.965854Z digest=sha256:dd26ac8f2f5a279d215f2230be938d0f472fdb817ebd2df2dbb4c92590790a59

Observation 259d49b3-f1ee-41eb-aaf1-d1695b76edfe · outbound

This paper cites Eclipse: Efficient long-range video retrieval using sight and sound.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Eclipse: Efficient long-range video retrieval using sight and sound

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:06:06.123020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T15:06:04.970155Z digest=sha256:0a63b19f108999ba8710db832830bad24f1d15dc6d5ec5755484dc1bf051fcfb

Observation 508faf4f-05af-4693-85a3-1d318c14ad52 · outbound

This paper cites Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:06:06.108611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T15:06:04.974651Z digest=sha256:202b2c3deddcb37dd52b66749b225f2144e42bf662d87b215cdf4c25e7d8d7e1

Observation 4a015e24-5651-49eb-89cf-2299cc5cf88a · outbound

This paper cites Unified-IO: A Unified Model for Vision, Language, and Multi-Modal Tasks.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Unified-IO: A Unified Model for Vision, Language, and Multi-Modal Tasks

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T15:06:04.978924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:06:04.978924Z digest=sha256:5d12ec23bc5369fe031b7b26cc196b9c0bf245f70a9d61c4c8da09cf826d427c

Observation 2e37f6b3-5420-48b1-8ef4-4ef8f2d6b586 · outbound

This paper cites UniVL: A Unified Video and Language Pre-Training Model for Multimodal Understanding and Generation.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning UniVL: A Unified Video and Language Pre-Training Model for Multimodal Understanding and Generation

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-12T15:06:04.983697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:06:04.983697Z digest=sha256:94f1ec6f61e5fcf4b7859638a86b7196a1b170791a072976f9b2c30dd28598c9

Observation 544c7b6c-8643-477f-b6c9-ee214541a8b9 · outbound

This paper cites CLIP4Clip: An Empirical Study of CLIP for End to End Video Clip Retrieval.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning CLIP4Clip: An Empirical Study of CLIP for End to End Video Clip Retrieval

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T15:06:04.988340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:06:04.988340Z digest=sha256:5b8fd38b038c8c8b5852992506c256472563dfecaaef3bdfe59eab8c6c3fb5fb

Observation f99164ad-09ba-48b6-9d4f-52c978b8cdf9 · outbound

This paper cites Howto100m: Learning a text-video embedding by watching hundred million narrated video clips.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Howto100m: Learning a text-video embedding by watching hundred million narrated video clips

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:06:06.094585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T15:06:04.992820Z digest=sha256:454bb0a301efe4694cea1b9620b780942beb106c52c154bc9996a50c35ac42bd

Observation e7c291c4-96bf-4303-ba3f-4eaafe0b1ad6 · outbound

This paper cites Moments in time dataset: one million videos for event understanding.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Moments in time dataset: one million videos for event understanding

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:06:06.080134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T15:06:04.997164Z digest=sha256:10d0deb88aca28d281ff8b83c96aab770984e2da3ebf7bab5cd0528b773a908b

Observation 269d64be-8381-4c1f-bdfe-c09523415829 · outbound

This paper cites Re- thinking video vits: Sparse video tubes for joint image and video learning.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Re- thinking video vits: Sparse video tubes for joint image and video learning

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:06:06.066168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T15:06:05.001535Z digest=sha256:2058eab58ea92ae16cb598cc387a2c81fd482fcbab74275ea87d080cbfed8013

Observation 6df7cbe3-23f8-48ea-8e27-a3e6b204789d · outbound

This paper cites Dynamic pretraining of vision-language models.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Dynamic pretraining of vision-language models

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:06:06.050986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T15:06:05.005634Z digest=sha256:935c3d5fc1cdf4e986c67585a67dd324fcaf760f12e3f7482dee4f47baf4a911

Observation 4c1d06c5-315f-4441-b52c-d889c3de41f7 · outbound

This paper cites Mirasol3B: A multi- modal autoregressive model for time-aligned and contextual modalities.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Mirasol3B: A multi- modal autoregressive model for time-aligned and contextual modalities

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:06:06.035370Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T15:06:05.010128Z digest=sha256:835d48d7599523bd392742861016cc506ab41519b740581b5b368c63221d3806

Observation 70723b90-532a-4496-baa8-fe1dbfba2780 · outbound

This paper cites Timechat: A time-sensitive multimodallarge language model for long video understanding.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Timechat: A time-sensitive multimodallarge language model for long video understanding

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:06:06.020522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T15:06:05.014494Z digest=sha256:1b6dd6c321a62dbbf78b8de64a2cc1465d4b07de2662855f29a5d2d532761d52

Observation eecb69be-a250-46cb-a0e7-46e88aa337b4 · outbound

This paper cites Ryoo, AJ Piergiovanni, Anurag Arnab, Mostafa Dehghani, and Anelia Angelova.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Ryoo, AJ Piergiovanni, Anurag Arnab, Mostafa Dehghani, and Anelia Angelova

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:06:06.005351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T15:06:05.019001Z digest=sha256:8d21b629cea10de55036fac689b9803b5b740ae8d35bea63e239a91853cacabe

Observation 682ac631-1b37-4e47-9e54-d74537cbd27b · outbound

This paper cites Ryoo, Keerthana Gopalakrishnan, Kumara Ka- hatapitiya, Ted Xiao, Kanishka Rao, Austin Stone, Yao Lu, Julian Ibarz, and Anurag Arnab.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Ryoo, Keerthana Gopalakrishnan, Kumara Ka- hatapitiya, Ted Xiao, Kanishka Rao, Austin Stone, Yao Lu, Julian Ibarz, and Anurag Arnab

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:06:05.990104Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T15:06:05.023173Z digest=sha256:6e0770eb5a81de26a44b591dd1b650fb863502b47cce0c5b1df8d64f850bbdf9

Observation fe32cb33-f5a8-42cc-9296-cf38edf05b5b · outbound

This paper cites Tridet: Temporal action detection with relative boundary modeling.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Tridet: Temporal action detection with relative boundary modeling

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:06:05.975563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T15:06:05.027594Z digest=sha256:2b81a645d44b43faf2cae7fe37c0fe55f0fcb38eb4c30ae4fc695e2939e6889c

Observation 6873270e-a51b-4777-aa9c-5a4191ca1ef4 · outbound

This paper cites Flava: A foundational language and vision alignment model.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Flava: A foundational language and vision alignment model

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-12T15:06:05.032048Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:06:05.032048Z digest=sha256:2fd0dbff34c46e6dd30e6aa72fd26ab8e4cd2b61c4d8a89ce4c80d49d95314aa

Observation cc54361f-70fe-4445-813e-2264bf0e17bc · outbound

This paper cites Ucf101: A dataset of 101 human action classes from videos in the wild.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Ucf101: A dataset of 101 human action classes from videos in the wild

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:06:05.952057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T15:06:05.036535Z digest=sha256:0d35eaf023fd61f80a3b1b8919b5505b356249d72fc0cfd9419fcfe465c845fe

Observation 26e32ad5-2ac1-43f4-ad44-734d59ebf866 · outbound

This paper cites Long-form video-language pre- training with multimodal temporal contrastive learning.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Long-form video-language pre- training with multimodal temporal contrastive learning

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:06:05.937238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T15:06:05.040934Z digest=sha256:f051486592acc87cf544c2e0a4ce696836807381e462ee53b3deeffd38990edb

Observation 72a9ef8c-a6be-41f9-87f5-904215e960a3 · outbound

This paper cites Lxmert: Learning cross- modality encoder representations from transformers.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Lxmert: Learning cross- modality encoder representations from transformers

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:06:05.922326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T15:06:05.045382Z digest=sha256:ceca3ded6cadd6a455b5670e06ce359025bfc5caae69f4f6514c3290ef6926de

Observation eb9d6011-49f0-4598-a980-1aae07ec971d · outbound

This paper cites CLIP4Caption: CLIP for Video Caption.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning CLIP4Caption: CLIP for Video Caption

Reference 50

Resolution
verified exact
local_arxiv, observed 2026-08-12T15:06:05.353751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T15:06:05.049820Z digest=sha256:da09ffc58765fbb3583ce15bce0ef97ee7eace060d0af540f4b083a49513e5d0

Observation abaf2d17-1bd9-45b9-b3fa-65160eba14cb · outbound

This paper cites Cider: Consensus-based image description evalua- tion.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Cider: Consensus-based image description evalua- tion

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-12T15:06:05.054676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:06:05.054676Z digest=sha256:83c4a3550ec4be798c4f3b31b5dbc3f7e281a5ed05402404b11d24676239d2ff

Observation 84909b79-6f9a-438a-9985-b6e0e2125196 · outbound

This paper cites Bidirectional attentive fusion with context gating for dense video captioning.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Bidirectional attentive fusion with context gating for dense video captioning

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:06:05.897753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T15:06:05.058944Z digest=sha256:540eb6c57b6cace11056727938707f8ff3db5c7c0ce0a58fa19a6f2e46ceea36

Observation ecb8d9d6-4849-4fc7-a288-f1d99e3102e3 · outbound

This paper cites GIT: A Generative Image-to-text Transformer for Vision and Language.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning GIT: A Generative Image-to-text Transformer for Vision and Language

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-12T15:06:05.063232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:06:05.063232Z digest=sha256:f06d1cb930a4555d616eb210e21bf5af92c3328bc91df2fafdec6b0db9b89135

Observation 25f47987-40a1-46cb-8ca5-451b566e3853 · outbound

This paper cites Omnivid: A generative framework for universal video understanding.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Omnivid: A generative framework for universal video understanding

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:06:05.883020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T15:06:05.067933Z digest=sha256:c7c3271141614c56759ff918fa8479bd4ac1f4debaed4d7d26376207b1f74564

Observation ad699342-6b75-4888-a562-47e8d54a4a59 · outbound

This paper cites OFA: Unifying Architectures, Tasks, and Modalities Through a Simple Sequence-to-Sequence Learning Framework.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning OFA: Unifying Architectures, Tasks, and Modalities Through a Simple Sequence-to-Sequence Learning Framework

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-12T15:06:05.072176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:06:05.072176Z digest=sha256:d621b6d78f40538a21d84d06dc0939edea4e84267c5d40c69bc40d291f380c9c

Observation 2844a8a3-3fc3-485b-b257-5efe1d665c05 · outbound

This paper cites End-to-end dense video captioning with parallel decoding.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning End-to-end dense video captioning with parallel decoding

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:06:05.867617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T15:06:05.076799Z digest=sha256:05bb2cf6353f0bc404082fbf3413ef4edece0af79924bfcaaaa84037941dc0d8

Observation 3866e6e4-4cd8-4a15-8fd0-c16be4a0a892 · outbound

This paper cites Image as a Foreign Language: BEiT Pretraining for All Vision and Vision-Language Tasks.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Image as a Foreign Language: BEiT Pretraining for All Vision and Vision-Language Tasks

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-12T15:06:05.081141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:06:05.081141Z digest=sha256:91433a941b1cc68c0e7fa8d29fd8efea1036b6a39da89b7b36c317187962be0d

Observation 77202398-63d6-4180-bcfc-77319f3070f1 · outbound

This paper cites InternVideo: General Video Foundation Models via Generative and Discriminative Learning.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-12T15:06:05.085716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:06:05.085716Z digest=sha256:d77bdb45759dc0c4465fa5d77ee8bf41cd9a69a5393f6a26ab252644727b67c0

Observation 7c79ed25-29e0-4b4f-a0ea-a4ffcf36494a · outbound

This paper cites SimVLM: Simple Visual Language Model Pretraining with Weak Supervision.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning SimVLM: Simple Visual Language Model Pretraining with Weak Supervision

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-12T15:06:05.090407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:06:05.090407Z digest=sha256:190577353d67526c1f6fdea8f7230a6f49dfdd266b985dcf3e64080db8fa3635

Observation dedb013e-e3c7-44ee-9654-c3856e996084 · outbound

This paper cites Vl-bert: Pre-training of generic visual- linguistic representations.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Vl-bert: Pre-training of generic visual- linguistic representations

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:06:05.852537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T15:06:05.095077Z digest=sha256:6bc7eff6ab4161d31fe49180d1dafe8782ee4396dd78b0b0f8743c9412e86b11

Observation 3a2d0625-250b-4b1d-a5ab-ee791bbe4902 · outbound

This paper cites Towards long-form video understanding.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Towards long-form video understanding

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:06:05.838141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T15:06:05.099227Z digest=sha256:c288b71b27eed9fcd0eda4f94666a7045eb4b012161947653c2745e624c276d5

Observation 0ffc5d5d-b2d5-41a5-abc9-b57b2ba4d9fb · outbound

This paper cites Dibs: Enhanc- ing dense video captioning with unlabeled videos via pseudo boundary enrichment and online refinement.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Dibs: Enhanc- ing dense video captioning with unlabeled videos via pseudo boundary enrichment and online refinement

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:06:05.823791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T15:06:05.103524Z digest=sha256:bb8619c341c877b5218c07d710e08174b3ec9a6abf230675fd52ced2cef1249c

Observation 484716e0-2545-429b-98de-c9ffdd87a54e · outbound

This paper cites mPLUG-2: A Modularized Multi-modal Foundation Model Across Text, Image and Video.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning mPLUG-2: A Modularized Multi-modal Foundation Model Across Text, Image and Video

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-12T15:06:05.108091Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:06:05.108091Z digest=sha256:471f934c68413caf18fa7c847a3bfa125dd2a7a6388b1991996a0517675db5ab

Observation f11a6134-c2e5-4735-9bcf-4d2ba70a9829 · outbound

This paper cites VideoCoCa: Video-Text Modeling with Zero-Shot Transfer from Contrastive Captioners.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning VideoCoCa: Video-Text Modeling with Zero-Shot Transfer from Contrastive Captioners

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-12T15:06:05.112688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:06:05.112688Z digest=sha256:1d0c6108af1d37108730fcc93dafe2bd14ab3231354f195c522944f809465f20

Observation 4ea51806-a878-4d44-a798-21e8c1530627 · outbound

This paper cites Vid2seq: Large-scale pretraining of a vi- sual language model for dense video captioning.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Vid2seq: Large-scale pretraining of a vi- sual language model for dense video captioning

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:06:05.808297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T15:06:05.117231Z digest=sha256:8bd0fdda98858a4faebc598a47d51541753c3f25683c5ba4c9b68e3a4882d968

Observation f72021f7-070e-4aaa-8326-44b87f8ca744 · outbound

This paper cites Coca: Contrastive captioners are image-text foundation models.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Coca: Contrastive captioners are image-text foundation models

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-12T15:06:05.121604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:06:05.121604Z digest=sha256:221b784d4666420e0d20f44c4b7092653a2d2574d485a5c0d987499d4a4b4814

Observation c5448c75-05e4-445e-9a98-1004fab46069 · outbound

This paper cites Hierarchical video-moment retrieval and step-captioning.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Hierarchical video-moment retrieval and step-captioning

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:06:05.783130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T15:06:05.126063Z digest=sha256:a83ebc7c48202bb0f3d589e77787c819b1f3f6b6061120b656f56c1d34b253b9

Observation 7cd6d66b-6d22-4bc3-8426-b728a86d9794 · outbound

This paper cites Mer- lot: Multimodal neural script knowledge models.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Mer- lot: Multimodal neural script knowledge models

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:06:05.768846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T15:06:05.130393Z digest=sha256:a59b304b6580eae09a2f735b7d7bd113995935d36343e273e7f9c079e5e8f8fb

Observation 3c8c9704-693c-4510-a8dd-0692858bcee5 · outbound

This paper cites Actionformer: Lo- calizing moments of actions with transformers.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Actionformer: Lo- calizing moments of actions with transformers

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:06:05.754096Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T15:06:05.134742Z digest=sha256:ace61ad9c0511a3136db31df7e286c7223eb071d105df936c7e53e52755d2895

Observation 3faccc33-6290-4dae-ad00-4adfd538343c · outbound

This paper cites VinVL: Revisiting Visual Representations in Vision-Language Models.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning VinVL: Revisiting Visual Representations in Vision-Language Models

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-12T15:06:05.139120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:06:05.139120Z digest=sha256:90273ef1ba72f57422bb52fab46eb1400cffcc82989e82818b860f6c8a696e9a

Observation 113ed232-ae2d-4a3c-bdd6-b7a0a89bc616 · outbound

This paper cites Unifying event detec- tion and captioning as sequence generation via pre-training.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Unifying event detec- tion and captioning as sequence generation via pre-training

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:06:05.738616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T15:06:05.143718Z digest=sha256:9199307d298e470bf785bfe251b088236e153bbdb28228fae17ae83cb209d23a

Observation 786bfcf7-53f3-4260-9b5d-ef716e770842 · outbound

This paper cites Open-ended long-form video question answering via hierarchical convolutional self-attention networks.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Open-ended long-form video question answering via hierarchical convolutional self-attention networks

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:06:05.722709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T15:06:05.148081Z digest=sha256:6a81078321f25ea969a786bd34f4b7a037fd67c2d577972f46c2e55284075148

Observation d6f167a0-1b80-45b5-8251-af197a0a4cbf · outbound

This paper cites Towards automatic learning of procedures from web instructional videos.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Towards automatic learning of procedures from web instructional videos

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:06:05.708283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T15:06:05.152827Z digest=sha256:bd88c3cfe513c1bb01917a8f2c2c2fae31ce96932bbfde518974a6eb8b8430e5

Observation 01ea6556-418e-491e-94a0-f39784588f12 · outbound

This paper cites End-to-end dense video captioning with masked transformer.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning End-to-end dense video captioning with masked transformer

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:06:05.693993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T15:06:05.157215Z digest=sha256:c5d0e7df55b5ac8dc2c5b064d803666d7ebe435731c13e85cd93f110c11c585d

Observation 4b0e473b-4958-4e01-be53-f7d53ff8bc8b · outbound

This paper cites Streaming dense video captioning.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Streaming dense video captioning

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:06:05.679149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T15:06:05.161886Z digest=sha256:41ddd9ad4f6d3e25576fda15c3c1aba0516077bde6db2930a754408bcd6eee9f

Observation 981ea771-ddc3-4beb-b52f-4a1c4a08eb2e · outbound

This paper cites Towards Understanding Sample Variance in Visually Grounded Language Generation: Evaluations and Observations.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Towards Understanding Sample Variance in Visually Grounded Language Generation: Evaluations and Observations

Reference 76

Resolution
verified exact
local_arxiv, observed 2026-08-12T15:06:05.215806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T15:06:05.166374Z digest=sha256:ba38a98791ccad162409e416556fa094d0da0227776a34bf1b4e3b56a2d8f5a3

Observation 6919d268-601d-4346-ab9b-df06ec865e03 · outbound

This paper cites Thapliyal, William Yang Wang, and Radu Soricut.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Thapliyal, William Yang Wang, and Radu Soricut

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:06:05.664229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T15:06:05.171159Z digest=sha256:c10f35c46644f4bc7ef0780e77865e9c914add63f595cf08a8627c86629eb467

Pith citing papers

No inbound Pith citation observations are available.