Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-16T11:53:01.295207Z
Paper Citation Record · LEDGER
As of 19 August 2026, this Paper Citation Record lists 47 of 47 outbound references and 0 inbound Pith citation observations for arXiv:2504.14432.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-16T11:53:01.295207Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
47 of 47 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation f304bddb-a710-4735-a1f8-5705df4b77c9 · outbound
ResNetVLLM -- Multi-modal Vision LLM for the Video Understanding Task Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6caaec32-146f-4585-9263-2beb1770d82c · outbound
ResNetVLLM -- Multi-modal Vision LLM for the Video Understanding Task LLaMA: Open and Efficient Foundation Language Models
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2a3d2a0a-338b-415c-8e6a-80765a7e319b · outbound
ResNetVLLM -- Multi-modal Vision LLM for the Video Understanding Task Stanford alpaca: An instruction-following llama model,
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 278e9e95-9dee-48e8-b39d-177465329327 · outbound
ResNetVLLM -- Multi-modal Vision LLM for the Video Understanding Task Gpt-4 technical report,
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 11e441ed-de59-4337-ad8a-b1d44cea5798 · outbound
ResNetVLLM -- Multi-modal Vision LLM for the Video Understanding Task Chatgpt,
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation ea945b28-4b8d-4a52-83bb-052b7efb4391 · outbound
ResNetVLLM -- Multi-modal Vision LLM for the Video Understanding Task Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality,
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 5138eca9-1058-4b8b-b10e-dc469b71be32 · outbound
ResNetVLLM -- Multi-modal Vision LLM for the Video Understanding Task Instructblip: Towards general-purpose vision- language models with instruction tuning,
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 58c22334-2aa7-4cf8-8cce-55eaece7e4bc · outbound
ResNetVLLM -- Multi-modal Vision LLM for the Video Understanding Task LLaMA-Adapter V2: Parameter-Efficient Visual Instruction Model
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 28b26283-eaf9-46c8-935a-18b9b181ab48 · outbound
ResNetVLLM -- Multi-modal Vision LLM for the Video Understanding Task Visual Instruction Tuning
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8851c7de-d104-4564-a291-f29e56dfc6d4 · outbound
ResNetVLLM -- Multi-modal Vision LLM for the Video Understanding Task MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9727e838-16b8-4ad5-b237-1788a8b561ff · outbound
ResNetVLLM -- Multi-modal Vision LLM for the Video Understanding Task Locality and compositionality in zero-shot learning,
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 69adc038-f660-44f9-b7f4-27ad234afac9 · outbound
ResNetVLLM -- Multi-modal Vision LLM for the Video Understanding Task Learning transferable visual models from natural language supervision,
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 53acd390-7574-4ba7-b50f-d3af457f750e · outbound
ResNetVLLM -- Multi-modal Vision LLM for the Video Understanding Task Valley: Video Assistant with Large Language model Enhanced abilitY
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 62ff7709-975a-40ee-8763-a6574ccff3b1 · outbound
ResNetVLLM -- Multi-modal Vision LLM for the Video Understanding Task Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 056694e6-9ce4-442f-9bfa-a5f452209657 · outbound
ResNetVLLM -- Multi-modal Vision LLM for the Video Understanding Task Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0a4cadb3-24cc-453a-8095-56a4ab63a266 · outbound
ResNetVLLM -- Multi-modal Vision LLM for the Video Understanding Task VideoChat: Chat-Centric Video Understanding
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 03ea62b9-90c1-4cb6-8c93-f6389d52ae93 · outbound
ResNetVLLM -- Multi-modal Vision LLM for the Video Understanding Task MovieChat: From Dense Token to Sparse Memory for Long Video Understanding
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f254d20a-cbfa-475c-891b-cab57b1d8224 · outbound
ResNetVLLM -- Multi-modal Vision LLM for the Video Understanding Task Learning spatiotemporal features with 3d convolutional networks,
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 3e539da7-a84f-433b-a24a-2d6638754125 · outbound
ResNetVLLM -- Multi-modal Vision LLM for the Video Understanding Task Quo vadis, action recognition? a new model and the kinetics dataset,
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 37d98bee-ba20-498c-9bd2-637a8cb93ba7 · outbound
ResNetVLLM -- Multi-modal Vision LLM for the Video Understanding Task Joint learning of attended zero-shot features and visual-semantic mapping,
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 20177b18-fcaa-47fa-9e5d-44a34325ea9c · outbound
ResNetVLLM -- Multi-modal Vision LLM for the Video Understanding Task A Review of Generalized Zero-Shot Learning Methods
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 89429e84-b8f8-4fef-98a3-303e0fb326eb · outbound
ResNetVLLM -- Multi-modal Vision LLM for the Video Understanding Task Learning a deep embedding model for zero-shot learning,
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 75935a25-512b-4e67-b893-7fbb54d1b0c1 · outbound
ResNetVLLM -- Multi-modal Vision LLM for the Video Understanding Task Video question answering via gradually refined attention over appearance and motion,
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bb6fc5cb-8e23-4633-ba85-5cd65e039f23 · outbound
ResNetVLLM -- Multi-modal Vision LLM for the Video Understanding Task Tgif-qa: Toward spatio- temporal reasoning in visual question answering,
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 7b1f0db8-c479-4411-be56-cef475623f71 · outbound
ResNetVLLM -- Multi-modal Vision LLM for the Video Understanding Task Activitynet- qa: A dataset for understanding complex web videos via question answering,
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7656832f-b8f0-424e-924a-f9683f859924 · outbound
ResNetVLLM -- Multi-modal Vision LLM for the Video Understanding Task Latent dirichlet allocation,
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation f78adb90-cac7-48a7-bdfb-27a7dd9b51ac · outbound
ResNetVLLM -- Multi-modal Vision LLM for the Video Understanding Task Efficient estimation of word representations in vector space,
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 93315135-3f80-4e4e-b595-389b4aede8e1 · outbound
ResNetVLLM -- Multi-modal Vision LLM for the Video Understanding Task Skip-thought vectors,
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 38ce0715-2bf0-48d1-bfe7-6d0469f846b7 · outbound
ResNetVLLM -- Multi-modal Vision LLM for the Video Understanding Task Distributed representations of sentences and documents,
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation ff98a2ae-8787-494f-a374-ab359de1a45e · outbound
ResNetVLLM -- Multi-modal Vision LLM for the Video Understanding Task A neural proba- bilistic language model,
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 0c37b4dd-802a-489e-a702-936c6219f618 · outbound
ResNetVLLM -- Multi-modal Vision LLM for the Video Understanding Task Language Models are Few-Shot Learners
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d05b9459-8a23-4a8a-bee2-7e4da5c94e42 · outbound
ResNetVLLM -- Multi-modal Vision LLM for the Video Understanding Task Energy Tank-Based Policies for Robust Aerial Physical Interaction with Moving Objects
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation d39ae4f1-3396-4ae7-b745-98df744e60bc · outbound
ResNetVLLM -- Multi-modal Vision LLM for the Video Understanding Task OPT: Open Pre-trained Transformer Language Models
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 70888efd-2d1a-42eb-b88f-26df1aec08d8 · outbound
ResNetVLLM -- Multi-modal Vision LLM for the Video Understanding Task Flamingo: A visual language model for few-shot learning,
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation b99605de-16e8-49a2-8595-5a15dac97fdf · outbound
ResNetVLLM -- Multi-modal Vision LLM for the Video Understanding Task Towards general purpose vision systems: An end-to-end task-agnostic vision-language architecture,
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 1318fc6c-a02a-4ad8-893f-e3974fc55ec3 · outbound
ResNetVLLM -- Multi-modal Vision LLM for the Video Understanding Task Class-agnostic object detection with multi-modal transformer,
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation ad11aa57-caf0-4881-8015-e5d4f645d8a4 · outbound
ResNetVLLM -- Multi-modal Vision LLM for the Video Understanding Task Bridg- ing the gap between object and image-level representations for open- vocabulary detection,
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation c466c31f-8346-4db8-8666-49e04d9a59b0 · outbound
ResNetVLLM -- Multi-modal Vision LLM for the Video Understanding Task Open-vocabulary semantic segmentation with mask-adapted clip,
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 83e32dad-0985-4a08-afed-7c6f3a8c4d07 · outbound
ResNetVLLM -- Multi-modal Vision LLM for the Video Understanding Task Language-grounded indoor 3d semantic segmentation in the wild,
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation bdd060f2-07d9-4739-9535-d14d05633eaf · outbound
ResNetVLLM -- Multi-modal Vision LLM for the Video Understanding Task Expanding language-image pretrained models for general video recognition,
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation ec746aca-0426-400e-9372-d412337a601a · outbound
ResNetVLLM -- Multi-modal Vision LLM for the Video Understanding Task ActionCLIP: A New Paradigm for Video Action Recognition
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a46ef225-0370-40f3-8308-8689c32b7d34 · outbound
ResNetVLLM -- Multi-modal Vision LLM for the Video Understanding Task Finetuned clip models are efficient video learners,
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 8417e37e-6722-40b1-9afb-9a3f9743a26e · outbound
ResNetVLLM -- Multi-modal Vision LLM for the Video Understanding Task Deep residual learning for image recognition,
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e0eebc19-9393-4a68-9cad-bba3fedf1478 · outbound
ResNetVLLM -- Multi-modal Vision LLM for the Video Understanding Task Temporal segment networks: Towards good practices for deep action recognition,
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0604e75d-9422-4794-ab3b-cc9e8cf058a7 · outbound
ResNetVLLM -- Multi-modal Vision LLM for the Video Understanding Task Activi- tynet: A large-scale video benchmark for human activity understanding,
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 0ba76b65-e53d-49a9-8a97-f5deaed3158c · outbound
ResNetVLLM -- Multi-modal Vision LLM for the Video Understanding Task BT-Adapter: Video Conversation is Feasible Without Video Instruction Tuning
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c0097f47-a2fb-420e-999e-701f5f012d1a · outbound
ResNetVLLM -- Multi-modal Vision LLM for the Video Understanding Task Zero-shot video question answering via frozen bidirectional language models,
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
No inbound Pith citation observations are available.