Pith. sign in

Paper Citation Record · LEDGER

Movie2Story: A framework for understanding videos and telling stories in the form of novel text

As of 13 August 2026, this Paper Citation Record lists 58 of 58 outbound references and 0 inbound Pith citation observations for arXiv:2412.14965.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.14965 v2

Coverage vector

measured 58 of 58 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T11:48:07.259130Z

measured 58 of 58 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

58 of 58 outbound references displayed

  • verified exact2
  • verified fuzzy0
  • unresolved56
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation cd9a2d0f-4f81-454b-86d7-ca7e8b50d817 · outbound

This paper cites Flamingo: a Visual Language Model for Few-Shot Learning.

Movie2Story: A framework for understanding videos and telling stories in the form of novel text Flamingo: a Visual Language Model for Few-Shot Learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T11:48:07.102344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:48:07.102344Z digest=sha256:67f6cd9609c16ba6ec21e12c7f222dcc370e025e094950515413789080d07a5a

Observation 6629e209-1133-4450-b9eb-37a1247dfd21 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Movie2Story: A framework for understanding videos and telling stories in the form of novel text Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T11:48:07.106481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:48:07.106481Z digest=sha256:e84903b7f31d25becc346a2d72501fa0f7a94ace25a297fcb7a6e65833857390

Observation 665c0f6b-3883-4577-8704-1dba32b7c9c1 · outbound

This paper cites HTS-AT: A Hierarchical Token-Semantic Audio Transformer for Sound Classification and Detection.

Movie2Story: A framework for understanding videos and telling stories in the form of novel text HTS-AT: A Hierarchical Token-Semantic Audio Transformer for Sound Classification and Detection

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T11:48:07.109873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:48:07.109873Z digest=sha256:f8896b37621a0884a537292726fa3312d87a8408cb979904dc3dc775c5b2aced

Observation 5ddfea5d-024b-46a4-b01a-fb5db9d6ef34 · outbound

This paper cites BEATs: Audio Pre-Training with Acoustic Tokenizers.

Movie2Story: A framework for understanding videos and telling stories in the form of novel text BEATs: Audio Pre-Training with Acoustic Tokenizers

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T11:48:07.113036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:48:07.113036Z digest=sha256:42b6fcd460f3f5e2f752b068c455649f4478a2099737042d2a9a7fe071a5542d

Observation b93bcddd-58ff-49db-b644-c6af165c04fe · outbound

This paper cites InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks.

Movie2Story: A framework for understanding videos and telling stories in the form of novel text InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T11:48:07.116192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:48:07.116192Z digest=sha256:886d747a9e7a2846382c0f7192fe397b1da15fda9e0d08528f24d4636a518cd2

Observation d7e4111d-1e26-4fea-a0c8-528596a46b7a · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

Movie2Story: A framework for understanding videos and telling stories in the form of novel text VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T11:48:07.119559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:48:07.119559Z digest=sha256:8c4a1de74af39c9e0d007dbf8e70d5e5ca3f4cc85130efd901f0496b3756cee5

Observation f96e29d8-7ca7-4892-82c4-dc466d4740d6 · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

Movie2Story: A framework for understanding videos and telling stories in the form of novel text BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T11:48:07.122823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:48:07.122823Z digest=sha256:fea3e470055425cc6b9835bf1d1b51b1d3c89f9a33071eddddbe87390b2316f8

Observation ef7a6140-d6ea-41cd-84e2-b95b63334213 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

Movie2Story: A framework for understanding videos and telling stories in the form of novel text An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T11:48:07.125832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:48:07.125832Z digest=sha256:a059d38dacb03fe2b964de6e6f0694b91a465eaa328d7ebc7909ebe5bee8a0f9

Observation de6d5932-b665-4bb6-92a6-5f3642cdbeba · outbound

This paper cites PaLM-E: An Embodied Multimodal Language Model.

Movie2Story: A framework for understanding videos and telling stories in the form of novel text PaLM-E: An Embodied Multimodal Language Model

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T11:48:07.128749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:48:07.128749Z digest=sha256:e2e8655de9fe1fdca5235349c5102ecc8c21d4ba24b77324ed4608711e460162

Observation 7ed0f38a-da12-4fa4-9290-714370f0ffd6 · outbound

This paper cites The Llama 3 Herd of Models.

Movie2Story: A framework for understanding videos and telling stories in the form of novel text The Llama 3 Herd of Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T11:48:07.131658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:48:07.131658Z digest=sha256:c4dfa0cea21953643edcc4e202461a2d13c627cc74eabccc666d764e5d0b67db

Observation e19bf70a-4876-4679-aec0-124eb466a99b · outbound

This paper cites Mistral 7B.

Movie2Story: A framework for understanding videos and telling stories in the form of novel text Mistral 7B

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T11:48:07.133826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:48:07.133826Z digest=sha256:9872a40a135b8ea40976322ace0b096d46bc753e169bcc607d01a22f9257553d

Observation 3f4e0708-b829-4151-9a3b-7db13f38056a · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

Movie2Story: A framework for understanding videos and telling stories in the form of novel text LoRA: Low-Rank Adaptation of Large Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T11:48:07.136034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:48:07.136034Z digest=sha256:62024ea84f2a36b86335edfde3a484030c64a58dee5878d11da7b1a6a921ddcc

Observation 2c77b2f8-6403-4a6f-8bcc-1d2915e1130e · outbound

This paper cites pyannote.audio: neural building blocks for speaker diarization.

Movie2Story: A framework for understanding videos and telling stories in the form of novel text pyannote.audio: neural building blocks for speaker diarization

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T11:48:07.138051Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:48:07.138051Z digest=sha256:5b61938116ec8ac0674c319e1e0d18ade17a6d400631db2dca183fdfa888c3f2

Observation 9533073b-bed7-41ba-8f87-3d69b11d8777 · outbound

This paper cites Qwen Technical Report.

Movie2Story: A framework for understanding videos and telling stories in the form of novel text Qwen Technical Report

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T11:48:07.140182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:48:07.140182Z digest=sha256:13775c0fcc120656dae5b27a7bf307616b3c2488070f7382041d09ae0e417b48

Observation 918f1167-b37f-4b81-88a2-be8fba7523a2 · outbound

This paper cites GPT-4o System Card.

Movie2Story: A framework for understanding videos and telling stories in the form of novel text GPT-4o System Card

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T11:48:07.142835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:48:07.142835Z digest=sha256:9817c922b2000e2690771a1755ef44d268433b31dc2cf7bbe839cb4920d86836

Observation e7f9a7a9-556d-458a-bfbb-362be4585dec · outbound

This paper cites emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation.

Movie2Story: A framework for understanding videos and telling stories in the form of novel text emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T11:48:07.145680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:48:07.145680Z digest=sha256:65589ee39766c78ed989938474589e92740ec3bac82710437d2ab10e0c1ad205

Observation 1587efce-0860-4235-a2f9-c4753eca11b1 · outbound

This paper cites an unresolved cited work.

Movie2Story: A framework for understanding videos and telling stories in the form of novel text Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-11T11:48:07.919248Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T11:48:07.148685Z digest=sha256:d0ab110bd3c94ba442cbfcb7e246711306bf556e8db2dbc9041ab084e9de406e

Observation 918c887c-7417-4ee4-8c50-3483b9cf65c3 · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

Movie2Story: A framework for understanding videos and telling stories in the form of novel text Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T11:48:07.152211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:48:07.152211Z digest=sha256:fba9555f6b848e728835812d8d449e3b6e824e65152ac7d4918f34e46aff5bc8

Observation 0d037024-1051-4a3c-a188-cd172d282873 · outbound

This paper cites VITA: Towards Open-Source Interactive Omni Multimodal LLM.

Movie2Story: A framework for understanding videos and telling stories in the form of novel text VITA: Towards Open-Source Interactive Omni Multimodal LLM

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T11:48:07.155335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:48:07.155335Z digest=sha256:eb13d1b34741c328cddaa0e341af408fb63aeadeb6b737ec34a772d470447cb4

Observation b37fe7cd-2232-484e-bdc8-a2fada3b1a5d · outbound

This paper cites ImageBind: One Embedding Space To Bind Them All.

Movie2Story: A framework for understanding videos and telling stories in the form of novel text ImageBind: One Embedding Space To Bind Them All

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T11:48:07.158316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:48:07.158316Z digest=sha256:c4f5967f0a8366406d8ab7cc8ece3b0a3e61e31ac9c836a8a82b24b97015556a

Observation 6f95db44-06c5-461b-b93d-2fb0f6537abb · outbound

This paper cites The "something something" video database for learning and evaluating visual common sense.

Movie2Story: A framework for understanding videos and telling stories in the form of novel text The "something something" video database for learning and evaluating visual common sense

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T11:48:07.161317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:48:07.161317Z digest=sha256:35bccf34c67d69be3afc9765d25be298ad9d6cdf9d4ccb4cabdbd84198ab1ee0

Observation 89da694d-c378-4495-bbaf-48d4a02e93df · outbound

This paper cites Language Is Not All You Need: Aligning Perception with Language Models.

Movie2Story: A framework for understanding videos and telling stories in the form of novel text Language Is Not All You Need: Aligning Perception with Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T11:48:07.164357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:48:07.164357Z digest=sha256:94b98459568f6f541835ec34117a1024d7e0cecf66b181e743193236c8e8fa3f

Observation 2ed91dd6-bca6-42cd-a65d-adfc38ce79bd · outbound

This paper cites The Kinetics Human Action Video Dataset.

Movie2Story: A framework for understanding videos and telling stories in the form of novel text The Kinetics Human Action Video Dataset

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T11:48:07.167213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:48:07.167213Z digest=sha256:8dba8d00eb0f8d05c4fa0477dc8ca86515608afe208b2d07076ca2ac4300a1fc

Observation 2507dc6f-705b-48cd-8927-98653dea6df3 · outbound

This paper cites an unresolved cited work.

Movie2Story: A framework for understanding videos and telling stories in the form of novel text Unresolved cited work

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T11:48:07.169934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:48:07.169934Z digest=sha256:8a69273f1e7cce9eb4f3bf7b6fdbd354b514326fd6cfd54ddd6c58397232b39b

Observation 57a20458-e932-4b76-93c1-4b894ad55745 · outbound

This paper cites NowYouSee Me: Context-Aware Automatic Audio Description.

Movie2Story: A framework for understanding videos and telling stories in the form of novel text NowYouSee Me: Context-Aware Automatic Audio Description

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T11:48:07.172477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:48:07.172477Z digest=sha256:20e830cedc117679eed8a96ba01444b10008aa48403de0d60aaa394548fc6db4

Observation 0d52e13b-9464-4109-8aff-0babfde8d622 · outbound

This paper cites SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension.

Movie2Story: A framework for understanding videos and telling stories in the form of novel text SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T11:48:07.175428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:48:07.175428Z digest=sha256:8f53555a9ad5e4ce0ff52879cc869a5b7eb68fc876643db9c791ebe8324581f3

Observation ddbc04f6-4c79-4f35-b983-679adfe59a15 · outbound

This paper cites Kankanhalli.

Movie2Story: A framework for understanding videos and telling stories in the form of novel text Kankanhalli

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T11:48:07.178368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:48:07.178368Z digest=sha256:800857ffab459fd4ec0c5ca23d3cf5ef5e48b1ff3b9fc14aa6aae8247cfcc781

Observation fa077860-7905-48e7-87c5-3891165760c8 · outbound

This paper cites VideoChat: Chat-Centric Video Understanding.

Movie2Story: A framework for understanding videos and telling stories in the form of novel text VideoChat: Chat-Centric Video Understanding

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T11:48:07.180993Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:48:07.180993Z digest=sha256:78c064292fd390fde107891d2dcf1b2f0a53cc1870ad440435ac8df9f2e9a95c

Observation e5eb1154-0f5e-4f3b-97a6-4c520c7c4bb5 · outbound

This paper cites MVBench: A Comprehensive Multi-modal Video Understanding Benchmark.

Movie2Story: A framework for understanding videos and telling stories in the form of novel text MVBench: A Comprehensive Multi-modal Video Understanding Benchmark

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T11:48:07.183821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:48:07.183821Z digest=sha256:d3c5aaaddcafa9eea78d89ef0973b72fc8b08784190b7947808b1ce1f2a76a23

Observation 234b8113-a46c-43d2-a947-c42230ea32d9 · outbound

This paper cites Video-LLaVA: Learning United Visual Representation by Alignment Before Projection.

Movie2Story: A framework for understanding videos and telling stories in the form of novel text Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T11:48:07.186559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:48:07.186559Z digest=sha256:cc3681a7d3b9f20537966ca9a2b75a1377ed9157c7a7d2f06becd968ba072f8c

Observation 8279db1b-2f6a-461b-b949-46621b2e2bb9 · outbound

This paper cites an unresolved cited work.

Movie2Story: A framework for understanding videos and telling stories in the form of novel text Unresolved cited work

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T11:48:07.189316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:48:07.189316Z digest=sha256:9eac4d7434e832543bad891c0030e1830d55fd01ff1821967fc9fec46f889853

Observation c7417c79-a758-4d7a-af3d-68c16309da6b · outbound

This paper cites Visual Instruction Tuning.

Movie2Story: A framework for understanding videos and telling stories in the form of novel text Visual Instruction Tuning

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T11:48:07.191264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:48:07.191264Z digest=sha256:370603ddcbf8f505e1aa8ec3a4e8974e5eeaf36348ac07c87f42f048c0671d11

Observation cef6858c-35bc-46f4-8f41-19f74e14c841 · outbound

This paper cites MMBench: Is Your Multi-modal Model an All-around Player?.

Movie2Story: A framework for understanding videos and telling stories in the form of novel text MMBench: Is Your Multi-modal Model an All-around Player?

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T11:48:07.193389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:48:07.193389Z digest=sha256:1ca973c9c7763dfde44408461c320a955be979ac896b869bfd706c4587b5d5d6

Observation 96e9d748-0816-492a-909a-bb559c1a83cb · outbound

This paper cites Macaw-LLM: Multi-Modal Language Modeling with Image, Audio, Video, and Text Integration.

Movie2Story: A framework for understanding videos and telling stories in the form of novel text Macaw-LLM: Multi-Modal Language Modeling with Image, Audio, Video, and Text Integration

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T11:48:07.195569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:48:07.195569Z digest=sha256:c01cc940a2dba64ac54d23aa7e7b9b8a9d08192577eb1d45aa7e767e07fd67f7

Observation 6ec3c927-8318-4f6e-9ceb-87ba9fa6ae6d · outbound

This paper cites Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models.

Movie2Story: A framework for understanding videos and telling stories in the form of novel text Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T11:48:07.197598Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:48:07.197598Z digest=sha256:d4b1912a194163c3ac31e5d8afdb904f411a6cfccc8d9e5298bfbb8917b3db50

Observation caed6f36-bebe-43b8-9854-1642dccd3482 · outbound

This paper cites an unresolved cited work.

Movie2Story: A framework for understanding videos and telling stories in the form of novel text Unresolved cited work

Reference 36

Resolution
unresolved
raw_fallback, observed 2026-08-11T11:48:07.904681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T11:48:07.199634Z digest=sha256:203709a1c6b97795cd36b8b0290f10a22235ce76fe9ecb02f1b007924f1acfaf

Observation c9274df6-1dd0-4f2f-ab4c-c2e9eca1e918 · outbound

This paper cites an unresolved cited work.

Movie2Story: A framework for understanding videos and telling stories in the form of novel text Unresolved cited work

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T11:48:07.201536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:48:07.201536Z digest=sha256:17c96da6d8334aa31bea12c0d7f8b71606aceb404b94bbd8aca814cdff862ff3

Observation a9b6ac0b-271d-4707-9b02-efa9b456f11e · outbound

This paper cites Perception Test: A Diagnostic Benchmark for Multimodal Video Models.

Movie2Story: A framework for understanding videos and telling stories in the form of novel text Perception Test: A Diagnostic Benchmark for Multimodal Video Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-11T11:48:07.204032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:48:07.204032Z digest=sha256:82f7f80d1f887d388995096e68aab26ce4d8fb9c75e0379ffd8b464025822d6f

Observation ea0ada39-3cd6-43c0-a629-3b7cd5c75c09 · outbound

This paper cites Robust Speech Recognition via Large-Scale Weak Supervision.

Movie2Story: A framework for understanding videos and telling stories in the form of novel text Robust Speech Recognition via Large-Scale Weak Supervision

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T11:48:07.207064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:48:07.207064Z digest=sha256:c36742db7d777cea2495f5c91fb8c05da390de6cbf19e881160cc4ab0fccdc9c

Observation d6ecd06f-889b-4095-9958-d7e4bc40c222 · outbound

This paper cites an unresolved cited work.

Movie2Story: A framework for understanding videos and telling stories in the form of novel text Unresolved cited work

Reference 40

Resolution
verified exact
arxiv_id_nonexistent, observed 2026-08-11T11:48:07.471391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T11:48:07.209868Z digest=sha256:f3dbc2231477804f593c4f415b69790de54f6cd6ec1b60abd99f7c2dd2129e6e

Observation 556e9dc9-89ad-4f62-bd1a-07b39e84537f · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Movie2Story: A framework for understanding videos and telling stories in the form of novel text LLaMA: Open and Efficient Foundation Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T11:48:07.212367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:48:07.212367Z digest=sha256:fa87168881377af8b355db675fddc5a0ad4593c6dfc265c42d697717e3f810dd

Observation 7d677305-df6c-43d0-a5ed-c76897f128eb · outbound

This paper cites AV-SUPERB: A Multi-Task Evaluation Benchmark for Audio-Visual Representation Models.

Movie2Story: A framework for understanding videos and telling stories in the form of novel text AV-SUPERB: A Multi-Task Evaluation Benchmark for Audio-Visual Representation Models

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-08-11T11:48:07.367997Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T11:48:07.215478Z digest=sha256:b27bad06b6fc77c99f9f1e09899412b09730b1a7965aa5c999ba4a2a1502d1fc

Observation 9cfbe0f1-63d6-407a-8981-1ef078e0e4a7 · outbound

This paper cites InternVideo2: Scaling Foundation Models for Multimodal Video Understanding.

Movie2Story: A framework for understanding videos and telling stories in the form of novel text InternVideo2: Scaling Foundation Models for Multimodal Video Understanding

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T11:48:07.218794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:48:07.218794Z digest=sha256:1a5e38a03e0e626735d6c66db98b1bf002da0b9fb696bd26ea83d09410be051c

Observation 0ce6ac81-ac32-4b5c-ac18-878f9ed0616d · outbound

This paper cites InternVideo: General Video Foundation Models via Generative and Discriminative Learning.

Movie2Story: A framework for understanding videos and telling stories in the form of novel text InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-11T11:48:07.222604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:48:07.222604Z digest=sha256:f10c6b11425dc56daad649ad3b1fb7d1ef2cd2c46482255e5b9b53be12dac24a

Observation a182f44f-cc8e-4c4e-b4c9-72f1d94c21d1 · outbound

This paper cites NExT-QA:Next Phase of Question-Answering to Explaining Temporal Actions.

Movie2Story: A framework for understanding videos and telling stories in the form of novel text NExT-QA:Next Phase of Question-Answering to Explaining Temporal Actions

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-11T11:48:07.225318Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:48:07.225318Z digest=sha256:3ebd6aff615e72d158d666e4e530e62df3f4f61ec3805a8a0039ad22b5cc41c4

Observation cebcaa12-1d22-47e2-9e6b-74f8d28985bb · outbound

This paper cites FunQA: Towards Surprising Video Comprehension.

Movie2Story: A framework for understanding videos and telling stories in the form of novel text FunQA: Towards Surprising Video Comprehension

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-11T11:48:07.228234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:48:07.228234Z digest=sha256:0c7fd837faa6509404a94c423420135ec650c62c609d0e7a2f69fa01bed38726

Observation 22177569-1448-45b4-a04e-406c37531f4a · outbound

This paper cites an unresolved cited work.

Movie2Story: A framework for understanding videos and telling stories in the form of novel text Unresolved cited work

Reference 47

Resolution
unresolved
raw_fallback, observed 2026-08-11T11:48:07.895738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T11:48:07.231041Z digest=sha256:a49bd170c6300dcaf53ee0ae271f940e71298a0b179cf196f3ccdb9607ebd54d

Observation 543570b2-49ae-4006-8152-d932ba4c8155 · outbound

This paper cites an unresolved cited work.

Movie2Story: A framework for understanding videos and telling stories in the form of novel text Unresolved cited work

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-11T11:48:07.233568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:48:07.233568Z digest=sha256:3ebd4fb4d543f18a2fcd699d4b828837fa5721801242b0aa96aafe7c9090cab2

Observation 0e2f40d3-0671-4b39-95d9-7cfe102ba48b · outbound

This paper cites LVLM-eHub: A Comprehensive Evaluation Benchmark for Large Vision-Language Models.

Movie2Story: A framework for understanding videos and telling stories in the form of novel text LVLM-eHub: A Comprehensive Evaluation Benchmark for Large Vision-Language Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-11T11:48:07.236231Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:48:07.236231Z digest=sha256:b7680e6a5d76ef76d2d3488b0747fd543554577095cc8a04fb4b1c3932fc3561

Observation 0db9eb1f-fce8-40db-ab3f-3c67398d39fb · outbound

This paper cites Synchronized Video Storytelling: Generating Video Narrations with Structured Storyline.

Movie2Story: A framework for understanding videos and telling stories in the form of novel text Synchronized Video Storytelling: Generating Video Narrations with Structured Storyline

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-11T11:48:07.238929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:48:07.238929Z digest=sha256:b2e876cac34d35d8ba4ac5aa2d1994c8dd5c6f5341d74a9589a4c6d81f8541d4

Observation ee578662-6381-4d09-a55b-038b0767eb72 · outbound

This paper cites What is YOLOv8: An In-Depth Exploration of the Internal Features of the Next-Generation Object Detector.

Movie2Story: A framework for understanding videos and telling stories in the form of novel text What is YOLOv8: An In-Depth Exploration of the Internal Features of the Next-Generation Object Detector

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-11T11:48:07.241651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:48:07.241651Z digest=sha256:9fe7682030b8b83756c05b388f300c7a5bd574b27d1e86c935441997c066f207

Observation 3856e7d6-087c-4f6f-99e5-df0dc02fc371 · outbound

This paper cites mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality.

Movie2Story: A framework for understanding videos and telling stories in the form of novel text mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-11T11:48:07.244506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:48:07.244506Z digest=sha256:5d465e99fff6cdaaaef6322a3fc199fe8c2760a50b32ed17b3e5f70abb27c3f8

Observation 260020f6-a8cf-48d9-98b7-faa274e780c4 · outbound

This paper cites an unresolved cited work.

Movie2Story: A framework for understanding videos and telling stories in the form of novel text Unresolved cited work

Reference 53

Resolution
unresolved
raw_fallback, observed 2026-08-11T11:48:07.887203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T11:48:07.247202Z digest=sha256:f6c3a94a17199989e4f11d96c934fc8de86b783f67d63c3aeebf94d9db9000fd

Observation 5325b058-cfc7-444b-8ad9-06d1bc24f3c1 · outbound

This paper cites MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities.

Movie2Story: A framework for understanding videos and telling stories in the form of novel text MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-11T11:48:07.249685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:48:07.249685Z digest=sha256:cf48464affb5b8a960da5fe36ac5463d58ca113c66c6f1871893788ef278775a

Observation 98f1a2a0-8a7a-4c6e-8f91-c50e1ee5aad4 · outbound

This paper cites Distilling Vision-Language Models on Millions of Videos.

Movie2Story: A framework for understanding videos and telling stories in the form of novel text Distilling Vision-Language Models on Millions of Videos

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-11T11:48:07.251910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:48:07.251910Z digest=sha256:89aafc79bf5d657a01c5a34fc824907a3c1951f69fd6ea111f3173f497ae2328

Observation 9afde8d9-d360-4f23-b161-c02b8ba82897 · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

Movie2Story: A framework for understanding videos and telling stories in the form of novel text MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-11T11:48:07.254058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:48:07.254058Z digest=sha256:44bc20a9dec67a0c0030389239d5404115c6242190701699290b051914640980

Observation e4705e8a-b71d-4bba-a84e-73e898a89bb1 · outbound

This paper cites online" 'onlinestring :=.

Movie2Story: A framework for understanding videos and telling stories in the form of novel text online" 'onlinestring :=

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-11T11:48:07.256287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:48:07.256287Z digest=sha256:21949f9b844b3d956e7f607b9db3c4e9f2773c80471e25ec41312514e458c926

Observation ce3b0de5-1f1a-4ee4-8842-fb92cebbe38d · outbound

This paper cites write newline.

Movie2Story: A framework for understanding videos and telling stories in the form of novel text write newline

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-11T11:48:07.259130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:48:07.259130Z digest=sha256:f8d4c939ca8e557c584caf0b1d27fee3e1f8d1b87eddba9566826f6f40863606

Pith citing papers

No inbound Pith citation observations are available.