Pith. sign in

Paper Citation Record · LEDGER

Towards Long Video Understanding via Fine-detailed Video Story Generation

As of 21 August 2026, this Paper Citation Record lists 100 of 108 outbound references and 0 inbound Pith citation observations for arXiv:2412.06182.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.06182 v2

Coverage vector

measured 100 of 108 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T19:59:13.619204Z

measured 100 of 100 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

100 of 108 outbound references displayed

  • verified exact0
  • verified fuzzy57
  • unresolved43
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b76aa785-a8f7-40b9-9d24-9b0847fe086a · outbound

This paper cites Intermediary-guided bidi- rectional spatial-temporal aggregation network for video-based visible- infrared person re-identification,.

Towards Long Video Understanding via Fine-detailed Video Story Generation Intermediary-guided bidi- rectional spatial-temporal aggregation network for video-based visible- infrared person re-identification,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T19:59:13.299656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:59:13.299656Z digest=sha256:aa0c2108cf424a5ba1f84b7fc9e34e35da7a6ea5d2bd65867246e42971c48d0f

Observation 03b84715-37a0-4f07-9597-baef3232868b · outbound

This paper cites Video moment re- trieval via comprehensive relation-aware network,.

Towards Long Video Understanding via Fine-detailed Video Story Generation Video moment re- trieval via comprehensive relation-aware network,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T19:59:13.304058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:59:13.304058Z digest=sha256:157791dca67cb3613809bade96191b42a9cda4cca9ebcaf9a6fa3998417852be

Observation 3c3e072d-7723-4ee6-b14c-14fd0680d94e · outbound

This paper cites Self-supervised adversarial video summarizer with context latent sequence learning,.

Towards Long Video Understanding via Fine-detailed Video Story Generation Self-supervised adversarial video summarizer with context latent sequence learning,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T19:59:13.308146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:59:13.308146Z digest=sha256:9fff49b7dedbadc4381663829e490a8951ba7e0cd976fad8cbcf442a96418390

Observation 2c88bb87-0543-4ad9-8de1-7925763f3a63 · outbound

This paper cites Complementarity- aware space learning for video-text retrieval,.

Towards Long Video Understanding via Fine-detailed Video Story Generation Complementarity- aware space learning for video-text retrieval,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T19:59:13.312043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:59:13.312043Z digest=sha256:60109be7c1dcde168b0f46d49c7a632c440985ada08a334591e898fefb681089

Observation 157f858f-1c4c-4644-9cf2-07453d74dfcb · outbound

This paper cites Moma-lrg: Language-refined graphs for multi-object multi-actor activity parsing,.

Towards Long Video Understanding via Fine-detailed Video Story Generation Moma-lrg: Language-refined graphs for multi-object multi-actor activity parsing,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T19:59:13.315626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:59:13.315626Z digest=sha256:88c1d2963dd18dbb147e098ef8da983992c17ba6b3e44577beb6cc05031f16f9

Observation 3c5484a6-9cf2-4b62-8e17-7f1bf7c1b273 · outbound

This paper cites Videoclip: Contrastive pre- training for zero-shot video-text understanding,.

Towards Long Video Understanding via Fine-detailed Video Story Generation Videoclip: Contrastive pre- training for zero-shot video-text understanding,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T19:59:13.319397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:59:13.319397Z digest=sha256:1bf2f29251822aaff9584cc671852f2f8e868e65e6959bc5d6b7e3c07fb34dad

Observation 6ddc86b7-04cf-4c20-a1b1-d33e8b7c07ea · outbound

This paper cites Graph convolutional module for temporal action localization in videos,.

Towards Long Video Understanding via Fine-detailed Video Story Generation Graph convolutional module for temporal action localization in videos,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T19:59:13.323074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:59:13.323074Z digest=sha256:36625bd537b5792e975900da9b5e28acf65611bc0d8fd16229e826eb8a2c4750

Observation 304fefc4-3b33-4afd-a4bb-1de82d73b17c · outbound

This paper cites Evcap: Element- aware video captioning,.

Towards Long Video Understanding via Fine-detailed Video Story Generation Evcap: Element- aware video captioning,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T19:59:13.326324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:59:13.326324Z digest=sha256:fdfaa94f885078dab43b6d97e7d31b14601cf3ed85e514356a724a5f349f9757

Observation 246b13d7-365c-4a74-bc4c-c0679c3f3802 · outbound

This paper cites Multi-granularity interaction and integration network for video question answering,.

Towards Long Video Understanding via Fine-detailed Video Story Generation Multi-granularity interaction and integration network for video question answering,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T19:59:13.329312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:59:13.329312Z digest=sha256:74103629c3bb2dad49f2ebd7fa984ce8ac9e56bc11dc891f2e0b6ba72439f091

Observation 106cebde-0324-45a0-808f-ff7265d0a44d · outbound

This paper cites Video question answering with semantic disentanglement and reasoning,.

Towards Long Video Understanding via Fine-detailed Video Story Generation Video question answering with semantic disentanglement and reasoning,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T19:59:13.332350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:59:13.332350Z digest=sha256:4fd7984f5476fb02216772e1b2855881ad817ed8535cfec4a49bb1bf2f6dc909

Observation 548d2b2a-2e8e-473e-85f3-74925b113941 · outbound

This paper cites Multilevel semantic interaction alignment for video–text cross-modal retrieval,.

Towards Long Video Understanding via Fine-detailed Video Story Generation Multilevel semantic interaction alignment for video–text cross-modal retrieval,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T19:59:13.335894Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:59:13.335894Z digest=sha256:bb3b26b2c779acf98770fef218992e058a392a43a385082785d86082753af43e

Observation 37202589-7b99-4879-8fac-b0d933076c10 · outbound

This paper cites VideoChat: Chat-Centric Video Understanding.

Towards Long Video Understanding via Fine-detailed Video Story Generation VideoChat: Chat-Centric Video Understanding

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T19:59:13.339068Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:59:13.339068Z digest=sha256:a578c139f93ac923a078db6f93cff4ac93f89b07654825230e6766083c3c2041

Observation bc909dc9-ba3a-4750-bbe9-39769e529306 · outbound

This paper cites Video-llama: An instruction-tuned audio-visual language model for video understanding,.

Towards Long Video Understanding via Fine-detailed Video Story Generation Video-llama: An instruction-tuned audio-visual language model for video understanding,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T19:59:13.342311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:59:13.342311Z digest=sha256:9554fb2d3a6b77e25de61c6abdc83b2d2762e501925d8c34956f87c3da93a32f

Observation f77b4e5e-7d18-4849-a028-b1c48563771f · outbound

This paper cites MovieChat: From Dense Token to Sparse Memory for Long Video Understanding.

Towards Long Video Understanding via Fine-detailed Video Story Generation MovieChat: From Dense Token to Sparse Memory for Long Video Understanding

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T19:59:13.345138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:59:13.345138Z digest=sha256:974167d224614a1f3ac998148a874085b7db6a3beb31d94be5dc0e3f073e8f4f

Observation 831d8942-a346-4f40-9a3b-25cb57a1004c · outbound

This paper cites Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models.

Towards Long Video Understanding via Fine-detailed Video Story Generation Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T19:59:13.348682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:59:13.348682Z digest=sha256:ab84476a15e098d7d9035509fee3bdd27aedd33ebc767b9e26129b4640668807

Observation 89fee407-103d-46cb-831e-8e19fb6d605f · outbound

This paper cites Language models with image descriptors are strong few-shot video-language learners,.

Towards Long Video Understanding via Fine-detailed Video Story Generation Language models with image descriptors are strong few-shot video-language learners,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T19:59:13.352109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:59:13.352109Z digest=sha256:541ec9ebd0562002df93611dfe1c6f5d17ae47ac0b04f14f4fc8445282f642e7

Observation 68e75ac0-42de-4cbc-8e57-69a8d3917f5a · outbound

This paper cites Compressed video action recognition with dual-stream and dual-modal transformer,.

Towards Long Video Understanding via Fine-detailed Video Story Generation Compressed video action recognition with dual-stream and dual-modal transformer,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T19:59:13.355176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:59:13.355176Z digest=sha256:330e1d2a72a6e4bb07a7af5df0a9132109ea4ce7e0cfc082c5090b7f13de7ffd

Observation ccab0104-9ecf-4506-8be1-da3f20434797 · outbound

This paper cites Dynamic spatial focus for efficient compressed video action recognition,.

Towards Long Video Understanding via Fine-detailed Video Story Generation Dynamic spatial focus for efficient compressed video action recognition,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T19:59:13.358243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:59:13.358243Z digest=sha256:f837e25574d59aa955013250e539e0644864e8f5331cfcabd39150e8ce368ac4

Observation 5f1d5e46-a24d-44ee-88fa-eee4f7a9d064 · outbound

This paper cites Alignment-guided temporal atten- tion for video action recognition,.

Towards Long Video Understanding via Fine-detailed Video Story Generation Alignment-guided temporal atten- tion for video action recognition,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T19:59:13.361520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:59:13.361520Z digest=sha256:7789dda1eddbb4b940e4c0eaaf4311781428261ecfc88658180b8c369602dbd5

Observation 787aab86-06d4-43a0-befd-c74f18c3aa3b · outbound

This paper cites Slowfast networks for video recognition,.

Towards Long Video Understanding via Fine-detailed Video Story Generation Slowfast networks for video recognition,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T19:59:13.364500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:59:13.364500Z digest=sha256:428f125e7b34eb01bd8bd3832232d3f8269129e7cb9593d0d39411bca3188ec0

Observation 2071a66f-0e0e-4850-9335-33bc33a806d5 · outbound

This paper cites Temporal distinct representation learning for action recognition,.

Towards Long Video Understanding via Fine-detailed Video Story Generation Temporal distinct representation learning for action recognition,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T19:59:13.367934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:59:13.367934Z digest=sha256:fa33936e0692f73ec5beab740cf3c499d103f8e3941ab624b1b4942923671562

Observation 3d99e65b-f375-4487-aafe-5d7a2e96741b · outbound

This paper cites Truncate-split-contrast: a framework for learning from mislabeled videos,.

Towards Long Video Understanding via Fine-detailed Video Story Generation Truncate-split-contrast: a framework for learning from mislabeled videos,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T19:59:13.371085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:59:13.371085Z digest=sha256:20dc04d37f5d99f9ccc88a743a9af65109b95a3fa94b341aa58275ffc2ac966c

Observation de6591e7-2ab0-41c1-a586-ed130cdf71c1 · outbound

This paper cites Reading-strategy inspired visual representation learning for text-to- video retrieval,.

Towards Long Video Understanding via Fine-detailed Video Story Generation Reading-strategy inspired visual representation learning for text-to- video retrieval,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T19:59:13.374125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:59:13.374125Z digest=sha256:c1a96f6dd9582f7c7e2e16b101137870505c1e5e72702f4768c5b72062723a80

Observation 903fe238-f9d2-46cb-89ec-7c40649e1aff · outbound

This paper cites Use what you have: Video retrieval using representations from collaborative experts,.

Towards Long Video Understanding via Fine-detailed Video Story Generation Use what you have: Video retrieval using representations from collaborative experts,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T19:59:13.377335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:59:13.377335Z digest=sha256:aeb77b7e1a95459bff2adbd8ebb2a4463ac9b7a397491ba0f2dd93c645dc2f65

Observation 881de220-95ef-413a-8ee8-773a53b9718a · outbound

This paper cites Dual encoding for zero-example video retrieval,.

Towards Long Video Understanding via Fine-detailed Video Story Generation Dual encoding for zero-example video retrieval,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T19:59:13.380823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:59:13.380823Z digest=sha256:34a5032b28c2ea9b6eb25cc52efae7f5a4180f2679e02818038155bfe846e4f6

Observation 61a2a580-296b-4df4-bf20-aff3036720a2 · outbound

This paper cites Locvtp: Video-text pre-training for temporal localization,.

Towards Long Video Understanding via Fine-detailed Video Story Generation Locvtp: Video-text pre-training for temporal localization,

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T19:59:13.383984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:59:13.383984Z digest=sha256:bb1cdec8b1566ddc186333a9d87ccf382fdacfcb60a2cba7dc0003a6d1d56c23

Observation adad5951-16a8-4bfb-83c9-20e48977a0a0 · outbound

This paper cites Breaking winner-takes-all: Iterative-winners-out networks for weakly supervised temporal action localization,.

Towards Long Video Understanding via Fine-detailed Video Story Generation Breaking winner-takes-all: Iterative-winners-out networks for weakly supervised temporal action localization,

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T19:59:13.387374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:59:13.387374Z digest=sha256:f0339b4db20ea3d3a2d2635b6a23b910faafab04cc3228a230f5012f65ecbdae

Observation 8c562c82-db62-4be2-bd76-c3c12052f130 · outbound

This paper cites Cross time-frequency transformer for temporal action localization,.

Towards Long Video Understanding via Fine-detailed Video Story Generation Cross time-frequency transformer for temporal action localization,

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T19:59:13.390767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:59:13.390767Z digest=sha256:bc7682dd147864bc82b9499fb32419a56127875f202c8bfef068a6905fbb52b4

Observation 1740bebf-8852-4dcf-87ce-0ac1d1bf3b23 · outbound

This paper cites Slow motion matters: A slow motion enhanced network for weakly supervised temporal action localization,.

Towards Long Video Understanding via Fine-detailed Video Story Generation Slow motion matters: A slow motion enhanced network for weakly supervised temporal action localization,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T19:59:13.394100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:59:13.394100Z digest=sha256:a0f48984b39335e98520fe41d8d66df44ec5247e30697dd42eea35b4856425bd

Observation ce42a0f6-b6d2-4169-a8d3-cbee63e055e4 · outbound

This paper cites Long-form video- language pre-training with multimodal temporal contrastive learning,.

Towards Long Video Understanding via Fine-detailed Video Story Generation Long-form video- language pre-training with multimodal temporal contrastive learning,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T19:59:13.397280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:59:13.397280Z digest=sha256:8e7ffbeb1a979275e1ae0ce2949652646c234128f1adb14fdc18f4df785fcc8e

Observation 4320506b-7c55-417b-ad3d-d9a5b657c202 · outbound

This paper cites VideoGraph: Recognizing Minutes-Long Human Activities in Videos.

Towards Long Video Understanding via Fine-detailed Video Story Generation VideoGraph: Recognizing Minutes-Long Human Activities in Videos

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T19:59:13.400389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:59:13.400389Z digest=sha256:c15c1f49eeaac30d5a4377beba28de944bf69219bb76237ffec78df3ef66e2d0

Observation be66552c-aac9-465d-a353-7201b08941d3 · outbound

This paper cites Supervoxel attention graphs for long-range video modeling,.

Towards Long Video Understanding via Fine-detailed Video Story Generation Supervoxel attention graphs for long-range video modeling,

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T19:59:13.404329Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:59:13.404329Z digest=sha256:11c0a2c0e34de9eb093435bce98bb28f837ea47c4cdc52c91872cf2b925c453d

Observation 0c772c00-a4df-49fe-a93a-02cc99cf828f · outbound

This paper cites Long movie clip classification with state-space video models,.

Towards Long Video Understanding via Fine-detailed Video Story Generation Long movie clip classification with state-space video models,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:59:14.539583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T19:59:13.407515Z digest=sha256:8c1e1fbb0bd60a620131772edc5c3dda3b8e842ae7ebe5e2f958d7883a019a8f

Observation be7e0557-c7f6-41d0-93a9-e7b5db8f2549 · outbound

This paper cites S4nd: Modeling images and videos as multidimensional signals with state spaces,.

Towards Long Video Understanding via Fine-detailed Video Story Generation S4nd: Modeling images and videos as multidimensional signals with state spaces,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:59:14.528911Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T19:59:13.410492Z digest=sha256:38afe331c501da66c7024f40602e061e7ea680b8805d984329222f0953c306cf

Observation 6a883042-e723-4ce8-8dba-bfb316fce0f6 · outbound

This paper cites Selective structured state-spaces for long-form video understanding,.

Towards Long Video Understanding via Fine-detailed Video Story Generation Selective structured state-spaces for long-form video understanding,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:59:14.518072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T19:59:13.413725Z digest=sha256:d0138cfc8b9c257bd96ba8df2afd8e12af6467ecbb38c74c479b0c1a0b04642d

Observation 3c55eb5c-84ae-440d-ba0d-78efe043e69b · outbound

This paper cites Efficiently modeling long sequences with structured state spaces,.

Towards Long Video Understanding via Fine-detailed Video Story Generation Efficiently modeling long sequences with structured state spaces,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:59:14.507046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T19:59:13.416851Z digest=sha256:9fd6bbbb1488fa3d2989c43c17a2809d859f07467afb3592a13024c0dbd0f355

Observation b707d111-080e-4719-841a-d0e10ae281ac · outbound

This paper cites Mgsampler: An explainable sampling strategy for video action recognition,.

Towards Long Video Understanding via Fine-detailed Video Story Generation Mgsampler: An explainable sampling strategy for video action recognition,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:59:14.495316Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T19:59:13.419845Z digest=sha256:2139c57776bd7a5748da1446293ca1ceda0edfd4113e5b5d5d809a79b9c07b4a

Observation ec1f12c2-12fd-4792-8e34-5910efa98bbd · outbound

This paper cites Adaframe: Adaptive frame selection for fast video recognition,.

Towards Long Video Understanding via Fine-detailed Video Story Generation Adaframe: Adaptive frame selection for fast video recognition,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:59:14.484590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T19:59:13.422737Z digest=sha256:c4c2bf2530b725df77c714cfa6fdf3e7d9543bc71de58a667b72a856d8d5a908

Observation 7e4b57c6-eef4-4842-883e-935ad5c79ac4 · outbound

This paper cites Localizing moments in long video via multimodal guidance,.

Towards Long Video Understanding via Fine-detailed Video Story Generation Localizing moments in long video via multimodal guidance,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:59:14.474385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T19:59:13.425583Z digest=sha256:3cf994bc703f2f77502d49a6aeb41d7e20f07839130b8f179dbd73d8a8aff9fb

Observation 00a45254-eb61-4d1d-bec5-d4f60dcc6e25 · outbound

This paper cites Mad: A scalable dataset for language grounding in videos from movie audio descriptions,.

Towards Long Video Understanding via Fine-detailed Video Story Generation Mad: A scalable dataset for language grounding in videos from movie audio descriptions,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:59:14.463310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T19:59:13.428711Z digest=sha256:49720ec1a87d9123e6b43517ddb3669b76aed24f2c535a7e49abffd15d1300f8

Observation 97ae6fcd-5ebe-4097-9584-57e18ee9b50e · outbound

This paper cites End-to-end learning of visual representations from uncurated instructional videos,.

Towards Long Video Understanding via Fine-detailed Video Story Generation End-to-end learning of visual representations from uncurated instructional videos,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:59:14.452622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T19:59:13.431718Z digest=sha256:0a46ce25632c79c1b4df333ac1aa9274de6be452f0bdd18f131438afbbbd6c43

Observation bf90c54e-8f4c-4c1e-81b9-45c03751ff83 · outbound

This paper cites Merlot: Multimodal neural script knowledge models,.

Towards Long Video Understanding via Fine-detailed Video Story Generation Merlot: Multimodal neural script knowledge models,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:59:14.441521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T19:59:13.434656Z digest=sha256:ecc186f542913371121f0eb686f9ca0d4fed74057913554b5ed121396ac7d9d1

Observation 8a745a5d-ad95-4236-9c48-2f872ed5b492 · outbound

This paper cites Scaling up vision-language pre-training for image captioning,.

Towards Long Video Understanding via Fine-detailed Video Story Generation Scaling up vision-language pre-training for image captioning,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:59:14.431524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T19:59:13.437890Z digest=sha256:774f493c20778775c6ff933ad29c21baa39d8671134e293fb354633b339165cc

Observation 6ded8614-ee68-4727-827a-9f20bc76284e · outbound

This paper cites Gbc: Guided alignment and adaptive boosting clip bridging vision and language for robust action recognition,.

Towards Long Video Understanding via Fine-detailed Video Story Generation Gbc: Guided alignment and adaptive boosting clip bridging vision and language for robust action recognition,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:59:14.420598Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T19:59:13.440965Z digest=sha256:2994b934485e37b0e880e7eec924df53ad7ea19284b4697836a546beb5841481

Observation 8170842f-9f6f-46c9-b9a9-151258bcb288 · outbound

This paper cites Unsu- pervised pre-training for temporal action localization tasks,.

Towards Long Video Understanding via Fine-detailed Video Story Generation Unsu- pervised pre-training for temporal action localization tasks,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:59:14.409107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T19:59:13.444160Z digest=sha256:04b952f9fa08a17291d17bbd7f34595d21d3062a4dcc1d3ab52db7844f4b932e

Observation 548090d7-eebb-48e5-8077-948385896a68 · outbound

This paper cites Clip4clip: An empirical study of CLIP for end to end video clip retrieval and captioning,.

Towards Long Video Understanding via Fine-detailed Video Story Generation Clip4clip: An empirical study of CLIP for end to end video clip retrieval and captioning,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:59:14.397729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T19:59:13.447486Z digest=sha256:1d885a5f174feb99046d02bcf87a1fc1212e1344090b32f59a44b1395d17fcb9

Observation 7d6a0576-d691-432a-9b02-92d48310de48 · outbound

This paper cites Human action recognition and prediction: A survey,.

Towards Long Video Understanding via Fine-detailed Video Story Generation Human action recognition and prediction: A survey,

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-11T19:59:13.450578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:59:13.450578Z digest=sha256:6e207b4a64555051d1dca2da3c2200ddf1959c0b6df907960460586e582ef0cc

Observation 986671f5-f4f9-4321-aa71-3a98f247ffa4 · outbound

This paper cites A survey on video-based human action recognition: recent updates, datasets, challenges, and applications,.

Towards Long Video Understanding via Fine-detailed Video Story Generation A survey on video-based human action recognition: recent updates, datasets, challenges, and applications,

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:59:14.379512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T19:59:13.453369Z digest=sha256:547c15c696f5a7d59bbe2e9ccd59867389b0874cd3189831c79c57063c864cd5

Observation 2081cbe2-ef2e-4b4b-80e2-26cf5f84a614 · outbound

This paper cites Lavender: Unifying video-language understanding as masked lan- guage modeling,.

Towards Long Video Understanding via Fine-detailed Video Story Generation Lavender: Unifying video-language understanding as masked lan- guage modeling,

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:59:14.368806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T19:59:13.456277Z digest=sha256:7f5236555cee21e88e0f943f69f11041fe5927e10ad59b1d4591a951a44788eb

Observation be59d67c-e8a0-4cc3-b078-0cf00140ee16 · outbound

This paper cites Videomae: Masked autoencoders are data-efficient learners for self-supervised video pre- training,.

Towards Long Video Understanding via Fine-detailed Video Story Generation Videomae: Masked autoencoders are data-efficient learners for self-supervised video pre- training,

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:59:14.358415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T19:59:13.459592Z digest=sha256:8a6240cf7f7cc08f0ba4889d2cc2cb353cb1ef26b589a350ce650d915eb324b7

Observation a6411d7e-81c2-46b2-9290-6155fde763eb · outbound

This paper cites Videomae v2: Scaling video masked autoencoders with dual masking,.

Towards Long Video Understanding via Fine-detailed Video Story Generation Videomae v2: Scaling video masked autoencoders with dual masking,

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:59:14.347902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T19:59:13.463104Z digest=sha256:a8f3582781da5f0a8c87dc804e3c4289eef331140903c717d76e939236374c58

Observation a7316983-e468-4d51-9030-b1592160bf1b · outbound

This paper cites VIOLET : End-to-End Video-Language Transformers with Masked Visual-token Modeling.

Towards Long Video Understanding via Fine-detailed Video Story Generation VIOLET : End-to-End Video-Language Transformers with Masked Visual-token Modeling

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-11T19:59:13.466341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:59:13.466341Z digest=sha256:2a282270e80f867204bc9dff6c792d6897b20fe22dc8d431838e0dfe9a5039e7

Observation ba31b6c5-0eb5-4468-a395-30e98376d61e · outbound

This paper cites All in one: Exploring unified video-language pre-training,.

Towards Long Video Understanding via Fine-detailed Video Story Generation All in one: Exploring unified video-language pre-training,

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:59:14.336103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T19:59:13.470172Z digest=sha256:53fb18067c8cbb0a6dae65cdd06d030e62bceb72f53d406de71b6fa3d65d33b9

Observation 69b91bbb-8d7a-4fc6-9e55-b0d424f983e9 · outbound

This paper cites InternVideo: General Video Foundation Models via Generative and Discriminative Learning.

Towards Long Video Understanding via Fine-detailed Video Story Generation InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-11T19:59:13.473248Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:59:13.473248Z digest=sha256:7f37676552fcaad3f3cbbc01f95d4d8f4e102ce4ee54f400ce7f218d65ab9eb3

Observation 03d1f286-a7fd-485c-9b5e-b3d8dbeb4c64 · outbound

This paper cites Language mod- els are few-shot learners,.

Towards Long Video Understanding via Fine-detailed Video Story Generation Language mod- els are few-shot learners,

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:59:14.325979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T19:59:13.476828Z digest=sha256:237f4490ef52b9002523a1f0e594e6e4d65e37ef8bbd854e7b07a5f3c6ed00a1

Observation 98458942-a666-484f-8f2d-3d651a394c1b · outbound

This paper cites GLM: general language model pretraining with autoregressive blank infilling,.

Towards Long Video Understanding via Fine-detailed Video Story Generation GLM: general language model pretraining with autoregressive blank infilling,

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:59:14.315557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T19:59:13.480137Z digest=sha256:a40eff8b8c101b8ea32712fa150b18eba513b15b3d54058d2cb265ab91ceac5d

Observation c8db19b9-f461-438d-9453-3f5b81894773 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Towards Long Video Understanding via Fine-detailed Video Story Generation LLaMA: Open and Efficient Foundation Language Models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-11T19:59:13.483476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:59:13.483476Z digest=sha256:cb110111774bae8a754ff7aaeec2629f67a5a033c67c0eafc7c96be7b687ae11

Observation b41f2851-7240-4be8-bf5a-73bca0b77d97 · outbound

This paper cites MIST : Multi-modal iterative spatial-temporal transformer for long-form video question answering,.

Towards Long Video Understanding via Fine-detailed Video Story Generation MIST : Multi-modal iterative spatial-temporal transformer for long-form video question answering,

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:59:14.303595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T19:59:13.487191Z digest=sha256:e7d104c2c64ed63383f14881caee5a3a9cb481e9bad80f67cc9040ca239894e4

Observation cd159143-971a-4beb-bada-fac3eb9a60c5 · outbound

This paper cites A joint sequence fusion model for video question answering and retrieval,.

Towards Long Video Understanding via Fine-detailed Video Story Generation A joint sequence fusion model for video question answering and retrieval,

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:59:14.291561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T19:59:13.490546Z digest=sha256:709e2ee9bc2e63c1a9dca3faffca7607938754b7be6a138c1a18b32bd88e0398

Observation d2387de6-656c-40f1-ab0f-432ba582b063 · outbound

This paper cites Video question answering via gradually refined attention over appear- ance and motion,.

Towards Long Video Understanding via Fine-detailed Video Story Generation Video question answering via gradually refined attention over appear- ance and motion,

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:59:14.279595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T19:59:13.493599Z digest=sha256:962fb028d4c3e4466f7eab465c59d5f8834439f289c2a03efd52d7ea9650f936

Observation bff493ec-b099-46b2-b64f-2e3d8539f224 · outbound

This paper cites Activitynet-qa: A dataset for understanding complex web videos via question answering,.

Towards Long Video Understanding via Fine-detailed Video Story Generation Activitynet-qa: A dataset for understanding complex web videos via question answering,

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:59:14.268566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T19:59:13.496737Z digest=sha256:7f98e1d2cc9d62b91e10d1bb8814044cb453d2df1bd08be159bbcc979845b1ac

Observation 522cb6ed-6d9f-4a78-8400-76f8bf6ab18e · outbound

This paper cites Swinbert: End-to-end transformers with sparse attention for video captioning,.

Towards Long Video Understanding via Fine-detailed Video Story Generation Swinbert: End-to-end transformers with sparse attention for video captioning,

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:59:14.256936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T19:59:13.499776Z digest=sha256:1138e407a6a6b36fa3514b7aa4e2c01c9e524034212da017e232a6b39838ec68

Observation 08ad18dd-0240-4738-ae53-f7bf796daec9 · outbound

This paper cites Vid2seq: Large-scale pretraining of a visual language model for dense video captioning,.

Towards Long Video Understanding via Fine-detailed Video Story Generation Vid2seq: Large-scale pretraining of a visual language model for dense video captioning,

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:59:14.245299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T19:59:13.502597Z digest=sha256:ea83d34d1ab4ded0d1ac860b7ee6343b112221f716153ae646ec3183265af691

Observation 7cadfdc3-52c4-4d92-8868-78f9cb26701b · outbound

This paper cites Dense- captioning events in videos,.

Towards Long Video Understanding via Fine-detailed Video Story Generation Dense- captioning events in videos,

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:59:14.233671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T19:59:13.506128Z digest=sha256:ea79f7673480649c572e49a80d57d92182ba6ada99d929cc523499a0a0720f7c

Observation df2e5754-c242-4eb4-9377-621b106bdf0f · outbound

This paper cites BLIP-2: Bootstrapping language- image pre-training with frozen image encoders and large language models,.

Towards Long Video Understanding via Fine-detailed Video Story Generation BLIP-2: Bootstrapping language- image pre-training with frozen image encoders and large language models,

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:59:14.221672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T19:59:13.509171Z digest=sha256:43ca02dc7c998460f3a9985e953f258cf861c5740320445e3e854078915cd530

Observation 0f9707f9-ad5b-47f5-9be7-fc6146ad7ac3 · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

Towards Long Video Understanding via Fine-detailed Video Story Generation MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-11T19:59:13.512195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:59:13.512195Z digest=sha256:c56ec08884bcab57d1bd377486741c34a49c1956d6d99698254b90f37a327dc8

Observation 447ac7d9-1ccd-48aa-8c0b-d5ab82159b1f · outbound

This paper cites InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation.

Towards Long Video Understanding via Fine-detailed Video Story Generation InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-11T19:59:13.516220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:59:13.516220Z digest=sha256:1182f21c3c49b729ce10908beb0416110032d116b68d88b4596e1307261b88b6

Observation ebac29eb-40df-406a-84f7-88fdeb1499aa · outbound

This paper cites Pyscenedetect,.

Towards Long Video Understanding via Fine-detailed Video Story Generation Pyscenedetect,

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:59:14.209964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T19:59:13.519600Z digest=sha256:48abfddc7e8f984ae465f3a0fcdb03cce69299db2f471e633bcb952bd05ba5d4

Observation 4b1ddd92-00f5-44b7-9f41-36cc95970907 · outbound

This paper cites Decord: An efficient video loader for deep learning,.

Towards Long Video Understanding via Fine-detailed Video Story Generation Decord: An efficient video loader for deep learning,

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:59:14.197450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T19:59:13.522639Z digest=sha256:2373143c23206b98ee987a124ca1f740c65330b84604b54e6066bdb787d1a0d8

Observation f11bf45d-6f12-488f-b4fd-81bd7de680d8 · outbound

This paper cites Learning transferable visual models from natural language supervision,.

Towards Long Video Understanding via Fine-detailed Video Story Generation Learning transferable visual models from natural language supervision,

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-11T19:59:13.525872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:59:13.525872Z digest=sha256:162decec77f48afefa9d5a254d29f780983c5bfc467206c4e778d9fd095c11f4

Observation d5849239-8def-49a6-a4c3-cd867a0ed0b8 · outbound

This paper cites DINO: DETR with improved denoising anchor boxes for end-to-end object detection,.

Towards Long Video Understanding via Fine-detailed Video Story Generation DINO: DETR with improved denoising anchor boxes for end-to-end object detection,

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:59:14.179339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T19:59:13.529027Z digest=sha256:5f6dd75afb358c8a2d8d2e83a6e19d2c5a43ff86c050fe8f43864501b9e53ba0

Observation ee5186e6-8829-4917-80a0-799e771344e1 · outbound

This paper cites Sentence-bert: Sentence embeddings using siamese bert-networks,.

Towards Long Video Understanding via Fine-detailed Video Story Generation Sentence-bert: Sentence embeddings using siamese bert-networks,

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:59:14.168721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T19:59:13.532056Z digest=sha256:beafc23e3d737535d9574dcbafacca238f6ee6519c3d85c2182bdc9c343c3c91

Observation 11655f67-fe3d-4ae2-a455-a70c149ca17b · outbound

This paper cites Partially relevant video retrieval,.

Towards Long Video Understanding via Fine-detailed Video Story Generation Partially relevant video retrieval,

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:59:14.158363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T19:59:13.535135Z digest=sha256:e228c5a5276aad2457da043549deaf0ef364c2d393ddde6b8618c11ac57eca16

Observation 5a466a5f-d901-4c80-88a4-acaad7416c99 · outbound

This paper cites Joint searching and grounding: Multi-granularity video content retrieval,.

Towards Long Video Understanding via Fine-detailed Video Story Generation Joint searching and grounding: Multi-granularity video content retrieval,

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:59:14.147009Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T19:59:13.538757Z digest=sha256:953d39ffd89016e7b0cc26d7e8623fa6e0dd58fbedcacd0d9c9d7b75b97bc27c

Observation bef96942-deb3-4b17-ad8e-b537412a17ba · outbound

This paper cites Hit: Hierar- chical transformer with momentum contrast for video-text retrieval,.

Towards Long Video Understanding via Fine-detailed Video Story Generation Hit: Hierar- chical transformer with momentum contrast for video-text retrieval,

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:59:14.136686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T19:59:13.542232Z digest=sha256:73806e7eeb216daaaca63d5794f7f5c72354d5660515f0898c11b1241bc3fc7b

Observation 619faa62-256d-4b58-b41f-2d23bc20fbff · outbound

This paper cites Multi-modal trans- former for video retrieval,.

Towards Long Video Understanding via Fine-detailed Video Story Generation Multi-modal trans- former for video retrieval,

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:59:14.126362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T19:59:13.545156Z digest=sha256:305074ef83ec36d0c5bf05afdf088b21d262b389e57070e5e0bced50740af2b5

Observation 818e37cf-5c0d-4737-b57a-81c6e1766cc0 · outbound

This paper cites Cross-modal and hierarchical modeling of video and text,.

Towards Long Video Understanding via Fine-detailed Video Story Generation Cross-modal and hierarchical modeling of video and text,

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:59:14.115696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T19:59:13.548293Z digest=sha256:69de350e9bba02265316bd6a1a29cd78ee59c9113113433ccc9fcd9f98003df2

Observation 05a59bbb-5df8-486d-91ba-046535002610 · outbound

This paper cites Eclipse: Efficient long- range video retrieval using sight and sound,.

Towards Long Video Understanding via Fine-detailed Video Story Generation Eclipse: Efficient long- range video retrieval using sight and sound,

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:59:14.105621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T19:59:13.551159Z digest=sha256:ad57fc72d1dec500e0c65d4e755837857688d37005d627924047463a85fab7d3

Observation 7750a39d-77a5-481a-9aa4-1712f209efa4 · outbound

This paper cites TALL: temporal activity localization via language query,.

Towards Long Video Understanding via Fine-detailed Video Story Generation TALL: temporal activity localization via language query,

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:59:14.095823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T19:59:13.554216Z digest=sha256:e884a4b821d8f55d9d3f8c5bbd060df5a9556a1e2548cc1ba8943cfc39ba9c6a

Observation f586d4ae-e6cc-47e5-8c52-ee0c3c2a2acb · outbound

This paper cites Hollywood in homes: Crowdsourcing data collection for activity understanding,.

Towards Long Video Understanding via Fine-detailed Video Story Generation Hollywood in homes: Crowdsourcing data collection for activity understanding,

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:59:14.085989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T19:59:13.557624Z digest=sha256:205e057101c47f7369eb5c7c7c535cbbac67da46473b788eee9399b6d6c331ba

Observation 2c271b89-00d5-48d8-ad34-618c428d11a8 · outbound

This paper cites MSR-VTT: A large video descrip- tion dataset for bridging video and language,.

Towards Long Video Understanding via Fine-detailed Video Story Generation MSR-VTT: A large video descrip- tion dataset for bridging video and language,

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:59:14.076134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T19:59:13.560460Z digest=sha256:8a9295014688eb926205a915739e6ad41769651ecd18f0fcfca775a2bec31db1

Observation f42f821b-3b50-472a-875e-c61b44477382 · outbound

This paper cites Question generation via overgenerating transformations and ranking,.

Towards Long Video Understanding via Fine-detailed Video Story Generation Question generation via overgenerating transformations and ranking,

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:59:14.065916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T19:59:13.563116Z digest=sha256:fc44441465e6d8683fa1e35e093ce6d13cdd8d112889ea4a5ad5164d8ac428d7

Observation b12f594c-664e-4e6a-b75c-8f8707ec5244 · outbound

This paper cites Egoschema: A diagnostic benchmark for very long-form video language understanding,.

Towards Long Video Understanding via Fine-detailed Video Story Generation Egoschema: A diagnostic benchmark for very long-form video language understanding,

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:59:14.056879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T19:59:13.566380Z digest=sha256:d52ab80451a6f5c9f407eb2f3f9d5174516d002f79f5ee809d15c104ba54c145

Observation 4a7c3694-66ff-4de1-bb2f-002d1f1b1d72 · outbound

This paper cites Ego4d: Around the world in 3, 000 hours of egocentric video,.

Towards Long Video Understanding via Fine-detailed Video Story Generation Ego4d: Around the world in 3, 000 hours of egocentric video,

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:59:14.047172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T19:59:13.569192Z digest=sha256:876358cc4748cb6904c534dfa7dc8e8240a4b66ac2d983f7ec5fb21c37c42c87

Observation 884ac27b-36fd-4ab8-9bcd-0063a955b8f8 · outbound

This paper cites Next-qa: Next phase of question-answering to explaining temporal actions,.

Towards Long Video Understanding via Fine-detailed Video Story Generation Next-qa: Next phase of question-answering to explaining temporal actions,

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:59:14.036272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T19:59:13.572112Z digest=sha256:36772cce6e734d0e5420350ddd08a8416e2cbd9f8ad6e744ba47ff04e6ae280f

Observation 22780f74-4a8e-43b9-8b2e-19825e63da4e · outbound

This paper cites Quo vadis, action recognition? a new model and the kinetics dataset,.

Towards Long Video Understanding via Fine-detailed Video Story Generation Quo vadis, action recognition? a new model and the kinetics dataset,

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:59:14.026777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T19:59:13.574991Z digest=sha256:8fc8ee16762eb5f2dcefa2175312e38ffbe4c7f7d8b5a8ad15108858e1f1e4fd

Observation de50226d-8903-4b14-bd41-818fd009af55 · outbound

This paper cites Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena.

Towards Long Video Understanding via Fine-detailed Video Story Generation Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-11T19:59:13.578033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:59:13.578033Z digest=sha256:ba6aa9464201a7b65f1231f9c00eea8bbe37fc8e00314d030a2e504f95844f69

Observation 3b77bdc2-774f-4302-b3fb-3a29b29d2bd5 · outbound

This paper cites Sharegpt: Share your wildest chatgpt conversations with one click,.

Towards Long Video Understanding via Fine-detailed Video Story Generation Sharegpt: Share your wildest chatgpt conversations with one click,

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:59:14.017304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T19:59:13.581465Z digest=sha256:247d0543dd305ec2202b1782a3c4ff8bebd71693454f8fda95f145e3550bf18b

Observation 4e65467a-9dc8-4212-b7e0-f05e6b940993 · outbound

This paper cites AnglE-optimized Text Embeddings.

Towards Long Video Understanding via Fine-detailed Video Story Generation AnglE-optimized Text Embeddings

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-11T19:59:13.585145Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:59:13.585145Z digest=sha256:99e164d12df3d4480216d02147d1b1f26508957d3d39a13fd28af7cceed100a5

Observation 7be46fc4-02ac-4d22-af5d-02c447adc113 · outbound

This paper cites Video corpus moment retrieval with contrastive learning,.

Towards Long Video Understanding via Fine-detailed Video Story Generation Video corpus moment retrieval with contrastive learning,

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:59:14.007278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T19:59:13.588390Z digest=sha256:3964e20177a1708966538b0b1cc364b1f15fc2f9b6efbd6ea4f7f380c2fb7593

Observation 8875e9a0-e7ed-4f2b-9067-a3ebc767d51c · outbound

This paper cites TVR: A large-scale dataset for video-subtitle moment retrieval,.

Towards Long Video Understanding via Fine-detailed Video Story Generation TVR: A large-scale dataset for video-subtitle moment retrieval,

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:59:13.997127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T19:59:13.591759Z digest=sha256:215c296bfd93d775465f1ec2c2668e9b5e408103910130f1e0a0b764b737f617

Observation 3aeba3ae-12b5-4ea8-96d3-233932ff55b5 · outbound

This paper cites Dual learning with dynamic knowledge distillation for partially relevant video retrieval,.

Towards Long Video Understanding via Fine-detailed Video Story Generation Dual learning with dynamic knowledge distillation for partially relevant video retrieval,

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:59:13.986868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T19:59:13.594774Z digest=sha256:20f0b8012a3479e49e692ccaa7f49892b89805c5ddea1f3ae245de8dd599d80d

Observation f92a385a-4186-4f2f-af3c-81ae0c0f3a94 · outbound

This paper cites Video-LLaVA: Learning United Visual Representation by Alignment Before Projection.

Towards Long Video Understanding via Fine-detailed Video Story Generation Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-11T19:59:13.597852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:59:13.597852Z digest=sha256:493d9c3f6e5f082a163354bdd49208f8ee2ae98d9e2fa9a737477530867ef947

Observation ab79c81a-42d1-4f97-aa46-be999033b701 · outbound

This paper cites Just ask: Learning to answer questions from millions of narrated videos,.

Towards Long Video Understanding via Fine-detailed Video Story Generation Just ask: Learning to answer questions from millions of narrated videos,

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:59:13.976799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T19:59:13.601448Z digest=sha256:ea76f0d056b55d6dcdfdec4de5b5102b91459f47bfb39e4e3fb1e20972e78993

Observation 2528ebbc-87f5-4143-8735-33589eb2ea56 · outbound

This paper cites MERLOT RESERVE: neural script knowledge through vision and language and sound,.

Towards Long Video Understanding via Fine-detailed Video Story Generation MERLOT RESERVE: neural script knowledge through vision and language and sound,

Reference 95

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:59:13.965769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T19:59:13.604459Z digest=sha256:6f0458b5547fb83cc567388e86c7850643d3c3c38634359767f9d47953b8d664

Observation cf7169bb-7386-4afe-a9d6-dd790f1767c7 · outbound

This paper cites Zero-shot video question answering via frozen bidirectional language models,.

Towards Long Video Understanding via Fine-detailed Video Story Generation Zero-shot video question answering via frozen bidirectional language models,

Reference 96

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:59:13.955780Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T19:59:13.607548Z digest=sha256:7aa634bf8614588f7c401d4a8a6a4989c33e2f3dacc95dae8b6e7e0f1ccb929f

Observation 921d8ba2-5ff9-4525-8263-44a3a5f22613 · outbound

This paper cites Hitea: Hierarchical temporal-aware video-language pre-training,.

Towards Long Video Understanding via Fine-detailed Video Story Generation Hitea: Hierarchical temporal-aware video-language pre-training,

Reference 97

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:59:13.944568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T19:59:13.610461Z digest=sha256:e9cd2cae66cd36ef7e0812dc42773af2e2d8f3e3ee519872ba3feabbe3350270

Observation 0e4515bd-6b0e-47f2-80a4-eff5f7432f4b · outbound

This paper cites Self-chained image-language model for video localization and question answering,.

Towards Long Video Understanding via Fine-detailed Video Story Generation Self-chained image-language model for video localization and question answering,

Reference 98

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:59:13.933048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T19:59:13.613053Z digest=sha256:225bd1e1e335fcf30debe846f3e071a8fdc63ff6c23021aed0e9e719ad4866c5

Observation 22431014-9daa-4649-902f-4214ea607eb1 · outbound

This paper cites Memory consolidation enables long-context video understanding,.

Towards Long Video Understanding via Fine-detailed Video Story Generation Memory consolidation enables long-context video understanding,

Reference 99

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:59:13.922071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T19:59:13.616348Z digest=sha256:e3e3d7bb84d60de0420547ba44da6fb0ff7846d1b00b99dc050c16bd763d9bda

Observation 8f5f7710-bbd0-4204-ab66-f4012eaf31af · outbound

This paper cites mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality.

Towards Long Video Understanding via Fine-detailed Video Story Generation mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-11T19:59:13.619204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:59:13.619204Z digest=sha256:7c650a2c65b2ecfea61318e511c72388ae4f3892712ac05b1be80be5ee8e7a10

Pith citing papers

No inbound Pith citation observations are available.