Pith. sign in

Paper Citation Record · LEDGER

Bridging Vision and Language: Modeling Causality and Temporality in Video Narratives

As of 18 August 2026, this Paper Citation Record lists 23 of 23 outbound references and 0 inbound Pith citation observations for arXiv:2412.10720.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.10720 v1

Coverage vector

measured 23 of 23 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T15:44:54.926584Z

measured 23 of 23 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

23 of 23 outbound references displayed

  • verified exact2
  • verified fuzzy2
  • unresolved18
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ee7622d1-5dfa-46c8-9bb8-a4d93a8bba3e · outbound

This paper cites Vid2seq: Large-scale pretraining of a visual language model for dense video captioning,.

Bridging Vision and Language: Modeling Causality and Temporality in Video Narratives Vid2seq: Large-scale pretraining of a visual language model for dense video captioning,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T15:44:54.753530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:44:54.753530Z digest=sha256:e8a90c750ea1cbfde329a7f0424c158ea26a1074a1cd60a66ec7449616067d42

Observation 7a630dca-16d1-4102-86e9-6628a724d1ca · outbound

This paper cites Triple sequence generativ e adversarial nets for unsupervised image captioning,.

Bridging Vision and Language: Modeling Causality and Temporality in Video Narratives Triple sequence generativ e adversarial nets for unsupervised image captioning,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T15:44:54.766203Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:44:54.766203Z digest=sha256:80f94631203a4b2baf96cbcf846b9a61ccd94a6eb5f63e62fd15e090fb4b7d5f

Observation 1f7f8107-574d-4ca9-9461-74136da45bf9 · outbound

This paper cites Sketch storytelling,.

Bridging Vision and Language: Modeling Causality and Temporality in Video Narratives Sketch storytelling,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T15:44:54.778107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:44:54.778107Z digest=sha256:9574bc9f3cb7ed69a9ddc753eaa3031af63adf0e07432d77c6ff2187ea95a31c

Observation 4b1fbd5f-e1eb-4d48-860b-c6adca70a371 · outbound

This paper cites NarrativeBridge: Enhancing Video Captioning with Causal-Temporal Narrative.

Bridging Vision and Language: Modeling Causality and Temporality in Video Narratives NarrativeBridge: Enhancing Video Captioning with Causal-Temporal Narrative

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T15:44:54.791036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:44:54.791036Z digest=sha256:82259b15a3756e84b1a7dbfbe22ddea7ca8f4c1f59cd86624977f92ba6a09389

Observation 1b0f9cdd-c751-4ab1-b6cc-6445be56d32c · outbound

This paper cites Rethinking Visual Dependency in Long-Context Reasoning for Large Vision-Language Models.

Bridging Vision and Language: Modeling Causality and Temporality in Video Narratives Rethinking Visual Dependency in Long-Context Reasoning for Large Vision-Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T15:44:54.799390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:44:54.799390Z digest=sha256:3ca73ca95e9daf8cdcb85e23943d07d5c46a217bf8fa015b0ae89297d71e71cf

Observation e0fda45c-3af6-47f4-b2b8-6419469532d0 · outbound

This paper cites Visual in-context le arning for large vision-language models,.

Bridging Vision and Language: Modeling Causality and Temporality in Video Narratives Visual in-context le arning for large vision-language models,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T15:44:54.803663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:44:54.803663Z digest=sha256:e3853bbec7c491324a025ea781e8a66941215d102ebb12c2ae1a0c577521361b

Observation 421e803c-170a-4b62-bd17-366913b0296d · outbound

This paper cites Are We on the Right Way for Evaluating Large Vision-Language Models?.

Bridging Vision and Language: Modeling Causality and Temporality in Video Narratives Are We on the Right Way for Evaluating Large Vision-Language Models?

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T15:44:54.809836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:44:54.809836Z digest=sha256:d25337ec5bf5df25064155a544cb604bbb3bb36e278d1542a47d68e309cf0992

Observation 4e582af8-2a10-421e-bb8c-385d906994b7 · outbound

This paper cites VisionLLM v2: An End-to-End Generalist Multimodal Large Language Model for Hundreds of Vision-Language Tasks.

Bridging Vision and Language: Modeling Causality and Temporality in Video Narratives VisionLLM v2: An End-to-End Generalist Multimodal Large Language Model for Hundreds of Vision-Language Tasks

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T15:44:54.831133Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:44:54.831133Z digest=sha256:7833baa37156a17d9c8d804edf923649b8286c8cf79ad6bb16b40b486b28c9da

Observation fc0a0539-7834-4bc9-850e-c2416b0f0f81 · outbound

This paper cites A su rvey on multimodal large language models,.

Bridging Vision and Language: Modeling Causality and Temporality in Video Narratives A su rvey on multimodal large language models,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:44:56.029034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T15:44:54.838345Z digest=sha256:da18d2fc4eb708b755e914f4bcf72e8fb2850e046fbeb430439a3bb538a20572

Observation 34f69918-805c-4277-9c18-291ad00a3b46 · outbound

This paper cites Multimodal event transformer for i mage-guided story ending generation,.

Bridging Vision and Language: Modeling Causality and Temporality in Video Narratives Multimodal event transformer for i mage-guided story ending generation,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T15:44:54.842939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:44:54.842939Z digest=sha256:032571520885a16af862838d1a99dbf3cb43da0279dd5cb9141a31932b69ebf7

Observation 81961021-753c-4c11-b087-7455143f4495 · outbound

This paper cites Fine-Tuning Large Vision-Language Models as Decision-Making Agents via Reinforcement Learning.

Bridging Vision and Language: Modeling Causality and Temporality in Video Narratives Fine-Tuning Large Vision-Language Models as Decision-Making Agents via Reinforcement Learning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T15:44:54.848987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:44:54.848987Z digest=sha256:2f59e005c3da2d8d7676f794d90bba1f4de6da32e54c67079aed1b6ddb812f66

Observation c9795afb-6b79-4556-a418-6a729a2bbd57 · outbound

This paper cites Style-aware contrastive learning for multi-style image captioning,.

Bridging Vision and Language: Modeling Causality and Temporality in Video Narratives Style-aware contrastive learning for multi-style image captioning,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T15:44:54.855522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:44:54.855522Z digest=sha256:94addf242603d0888b0327063d4df581a1213d4511d35483bc55e4faf6a8b1be

Observation 332e282d-e0f1-4489-8562-954c63e89838 · outbound

This paper cites Improving cross-modal alignment for text-guided image inpaint- ing,.

Bridging Vision and Language: Modeling Causality and Temporality in Video Narratives Improving cross-modal alignment for text-guided image inpaint- ing,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:44:55.981517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T15:44:54.861556Z digest=sha256:6eb03db877f45d5c4fd26353fa12e623b70e7d98ab29f54f963a15e67e09eb09

Observation 4f0ddb9d-d718-4827-bd03-1fc2844c8e8c · outbound

This paper cites InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output.

Bridging Vision and Language: Modeling Causality and Temporality in Video Narratives InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T15:44:54.869813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:44:54.869813Z digest=sha256:75fb01429acab6e9c9e644ee52642b234a4b456f2677d1c129da4118850fffe2

Observation 38fb7a67-f1d4-430b-9953-5a35a52b049c · outbound

This paper cites Thread of Thought Unraveling Chaotic Contexts.

Bridging Vision and Language: Modeling Causality and Temporality in Video Narratives Thread of Thought Unraveling Chaotic Contexts

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T15:44:54.875118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:44:54.875118Z digest=sha256:9e14bf82e4c3653bdd2a8b89e8d7de3456505fad38572cf657cd76c09ee142f5

Observation 3681400b-2e88-4dc2-9723-10bb43ceb208 · outbound

This paper cites Streaming dense video captioning,.

Bridging Vision and Language: Modeling Causality and Temporality in Video Narratives Streaming dense video captioning,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T15:44:54.883452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:44:54.883452Z digest=sha256:04187bf90550752d60fd394e309e86ae81bf2bf119b3cda7986059ddfc5e4844

Observation 46227d4c-2f77-4c3c-a2c2-594de139b5e4 · outbound

This paper cites Livecap: Live video captioning wit h sequential encoding network,.

Bridging Vision and Language: Modeling Causality and Temporality in Video Narratives Livecap: Live video captioning wit h sequential encoding network,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T15:44:54.888620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:44:54.888620Z digest=sha256:5cc5ec69c3c9bc54eb397b1cfb81ccbbeaf3fdbd580ac4e05e25212e7f1a6f34

Observation 48aa3127-6575-4f59-8bc3-bbb94c45feb0 · outbound

This paper cites Ret rieval enhanced zero-shot video captioning,.

Bridging Vision and Language: Modeling Causality and Temporality in Video Narratives Ret rieval enhanced zero-shot video captioning,

Reference 18

Resolution
verified exact
raw_fallback, observed 2026-08-11T15:44:55.563935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T15:44:54.893937Z digest=sha256:8f218eca45a4e1bf6eb709ac7767c62ffad7b9916699c701c72f7c2c87b4bdc0

Observation d7fb0076-51c6-4b82-8ee5-05e8022df916 · outbound

This paper cites Accura te and fast compressed video captioning,.

Bridging Vision and Language: Modeling Causality and Temporality in Video Narratives Accura te and fast compressed video captioning,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T15:44:54.911301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:44:54.911301Z digest=sha256:434aaaf6ab5378350896b9ecbbc2089d7493a733f59e057ba725789b254bb611

Observation a6c4e297-062b-4f96-8d31-97a676d7646b · outbound

This paper cites Available: https://doi.org/10.48550/ar Xiv.2405.07046.

Bridging Vision and Language: Modeling Causality and Temporality in Video Narratives Available: https://doi.org/10.48550/ar Xiv.2405.07046

Reference 20

Resolution
verified exact
doi, observed 2026-08-11T15:44:55.151058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T15:44:54.901405Z digest=sha256:9ac0829eb064136d74bbe3bb84811c40c73cea179c24f2816602a7d0d0733f21

Observation c9161954-b612-47f5-9f4e-5d78b07a56b0 · outbound

This paper cites Video Captioning with Aggregated Features Based on Dual Graphs and Gated Fusion.

Bridging Vision and Language: Modeling Causality and Temporality in Video Narratives Video Captioning with Aggregated Features Based on Dual Graphs and Gated Fusion

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T15:44:54.916642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:44:54.916642Z digest=sha256:29a9659df08c0da1cd06c534ff8b95ab298c877f0033edb7248bdae43db8c610

Observation 8d5e2da0-b6e7-48f7-8378-34e76cb4a432 · outbound

This paper cites Video Captioning with Aggregated Features Based on Dual Graphs and Gated Fusion.

Bridging Vision and Language: Modeling Causality and Temporality in Video Narratives Video Captioning with Aggregated Features Based on Dual Graphs and Gated Fusion

Reference 2023

Resolution
malformed identifier
local_arxiv, observed 2026-08-11T15:44:54.992595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T15:44:54.926584Z digest=sha256:e227bcf75bc0f556d24a0d5f3bbe8e5d59582fae0b701c1671df7e3c14782d08

Observation d1daa596-d7c4-4448-8d5a-9e83036f87b1 · outbound

This paper cites Are We on the Right Way for Evaluating Large Vision-Language Models?.

Bridging Vision and Language: Modeling Causality and Temporality in Video Narratives Are We on the Right Way for Evaluating Large Vision-Language Models?

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-11T15:44:54.816312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:44:54.816312Z digest=sha256:2ed2d389f496636921d1ccf7dbab2de37b5ea69778d1fc9941c28b94f619ebe7

Pith citing papers

No inbound Pith citation observations are available.