Pith. sign in

Paper Citation Record · LEDGER

ResNetVLLM -- Multi-modal Vision LLM for the Video Understanding Task

As of 19 August 2026, this Paper Citation Record lists 47 of 47 outbound references and 0 inbound Pith citation observations for arXiv:2504.14432.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.14432 v1

Coverage vector

measured 47 of 47 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T11:53:01.295207Z

measured 47 of 47 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

47 of 47 outbound references displayed

  • verified exact1
  • verified fuzzy26
  • unresolved19
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f304bddb-a710-4735-a1f8-5705df4b77c9 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

ResNetVLLM -- Multi-modal Vision LLM for the Video Understanding Task Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-16T11:53:01.097617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:53:01.097617Z digest=sha256:13ea14e9e307d5b2e03de2385706636531c25ca17186d125902f266f67df6bb8

Observation 6caaec32-146f-4585-9263-2beb1770d82c · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

ResNetVLLM -- Multi-modal Vision LLM for the Video Understanding Task LLaMA: Open and Efficient Foundation Language Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-16T11:53:01.102120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:53:01.102120Z digest=sha256:3ceb91c46f43eb00a0b6e3f16ee94ce6368b413e34b101ed06b9e0f77235a491

Observation 2a3d2a0a-338b-415c-8e6a-80765a7e319b · outbound

This paper cites Stanford alpaca: An instruction-following llama model,.

ResNetVLLM -- Multi-modal Vision LLM for the Video Understanding Task Stanford alpaca: An instruction-following llama model,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:53:01.980865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:53:01.106292Z digest=sha256:525353a7dd55a63bcd485a491e46d58800c142a33fb4a05d17e7547d1409a02a

Observation 278e9e95-9dee-48e8-b39d-177465329327 · outbound

This paper cites Gpt-4 technical report,.

ResNetVLLM -- Multi-modal Vision LLM for the Video Understanding Task Gpt-4 technical report,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:53:01.966491Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:53:01.110305Z digest=sha256:14e0bcbc88e30d783b3096357790057d78e40003d56a4114430a3d98a11e5128

Observation 11e441ed-de59-4337-ad8a-b1d44cea5798 · outbound

This paper cites Chatgpt,.

ResNetVLLM -- Multi-modal Vision LLM for the Video Understanding Task Chatgpt,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:53:01.952923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:53:01.114679Z digest=sha256:38dc8bc636eb8d9fbecf381e8114725ee007e54862e08cbaaa91991e4edb288d

Observation ea945b28-4b8d-4a52-83bb-052b7efb4391 · outbound

This paper cites Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality,.

ResNetVLLM -- Multi-modal Vision LLM for the Video Understanding Task Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:53:01.938007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:53:01.118813Z digest=sha256:97c6b9949ba2ccaf0a7b09f72a0a6c73f170b4b091769163415f2c7ed665f923

Observation 5138eca9-1058-4b8b-b10e-dc469b71be32 · outbound

This paper cites Instructblip: Towards general-purpose vision- language models with instruction tuning,.

ResNetVLLM -- Multi-modal Vision LLM for the Video Understanding Task Instructblip: Towards general-purpose vision- language models with instruction tuning,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:53:01.923310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:53:01.123276Z digest=sha256:3b8841e7542f5f1a1f31a5964e7c02717658fca2ca54b21aa6da404deacac104

Observation 58c22334-2aa7-4cf8-8cce-55eaece7e4bc · outbound

This paper cites LLaMA-Adapter V2: Parameter-Efficient Visual Instruction Model.

ResNetVLLM -- Multi-modal Vision LLM for the Video Understanding Task LLaMA-Adapter V2: Parameter-Efficient Visual Instruction Model

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-16T11:53:01.127204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:53:01.127204Z digest=sha256:38d08af529b52c879360ea2e0122634851e16bc39fcfec950b4cf0abbe0b16d4

Observation 28b26283-eaf9-46c8-935a-18b9b181ab48 · outbound

This paper cites Visual Instruction Tuning.

ResNetVLLM -- Multi-modal Vision LLM for the Video Understanding Task Visual Instruction Tuning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-16T11:53:01.131649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:53:01.131649Z digest=sha256:a7d592b2e7c3c5b683d6d8d9dbae8dd37c032dd6be561f961d93fcde1fad3fc6

Observation 8851c7de-d104-4564-a291-f29e56dfc6d4 · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

ResNetVLLM -- Multi-modal Vision LLM for the Video Understanding Task MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-16T11:53:01.135610Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:53:01.135610Z digest=sha256:a4817dd9397f864ab77340d87171d5f919094b53ef502c5ad631495484204964

Observation 9727e838-16b8-4ad5-b237-1788a8b561ff · outbound

This paper cites Locality and compositionality in zero-shot learning,.

ResNetVLLM -- Multi-modal Vision LLM for the Video Understanding Task Locality and compositionality in zero-shot learning,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:53:01.908050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:53:01.139693Z digest=sha256:e29bac2f701b2c39267fc3993cbfe9bbf28577fa90cda8f54c1f5cf7982bcd3d

Observation 69adc038-f660-44f9-b7f4-27ad234afac9 · outbound

This paper cites Learning transferable visual models from natural language supervision,.

ResNetVLLM -- Multi-modal Vision LLM for the Video Understanding Task Learning transferable visual models from natural language supervision,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:53:01.891547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:53:01.143511Z digest=sha256:fc3e94372805907f71ea7afd42d27fde63ba6660675f49ff2056cecb64557169

Observation 53acd390-7574-4ba7-b50f-d3af457f750e · outbound

This paper cites Valley: Video Assistant with Large Language model Enhanced abilitY.

ResNetVLLM -- Multi-modal Vision LLM for the Video Understanding Task Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-16T11:53:01.147624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:53:01.147624Z digest=sha256:f6f8b2db248dfade04aacca28a2dd6887dff689a4c15156f6fe72d1bcbf4dc4d

Observation 62ff7709-975a-40ee-8763-a6574ccff3b1 · outbound

This paper cites Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models.

ResNetVLLM -- Multi-modal Vision LLM for the Video Understanding Task Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-16T11:53:01.151813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:53:01.151813Z digest=sha256:4d1060eefb1b1da77a8adceb6583f299e6fe16c236967f784baa24c7d81ac3d8

Observation 056694e6-9ce4-442f-9bfa-a5f452209657 · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

ResNetVLLM -- Multi-modal Vision LLM for the Video Understanding Task Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-16T11:53:01.155987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:53:01.155987Z digest=sha256:c404ff7bbb05c09bd8da7d0ad328a2a2ab1808e2bd34d55a8a764fc3eb58b032

Observation 0a4cadb3-24cc-453a-8095-56a4ab63a266 · outbound

This paper cites VideoChat: Chat-Centric Video Understanding.

ResNetVLLM -- Multi-modal Vision LLM for the Video Understanding Task VideoChat: Chat-Centric Video Understanding

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-16T11:53:01.160345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:53:01.160345Z digest=sha256:fe1ce831731466b4c1f3a077cdef3182459ccdfea36464a80cae069766121b51

Observation 03ea62b9-90c1-4cb6-8c93-f6389d52ae93 · outbound

This paper cites MovieChat: From Dense Token to Sparse Memory for Long Video Understanding.

ResNetVLLM -- Multi-modal Vision LLM for the Video Understanding Task MovieChat: From Dense Token to Sparse Memory for Long Video Understanding

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-16T11:53:01.164655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:53:01.164655Z digest=sha256:e64e99882e9445f8ca5628fe0e47906722d755ab942d19c67c9f390f229ffc44

Observation f254d20a-cbfa-475c-891b-cab57b1d8224 · outbound

This paper cites Learning spatiotemporal features with 3d convolutional networks,.

ResNetVLLM -- Multi-modal Vision LLM for the Video Understanding Task Learning spatiotemporal features with 3d convolutional networks,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:53:01.875686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:53:01.169251Z digest=sha256:a800425055de5672585bf3af8900ef63676ca6be7c5208ab0f820aecca30cf62

Observation 3e539da7-a84f-433b-a24a-2d6638754125 · outbound

This paper cites Quo vadis, action recognition? a new model and the kinetics dataset,.

ResNetVLLM -- Multi-modal Vision LLM for the Video Understanding Task Quo vadis, action recognition? a new model and the kinetics dataset,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-16T11:53:01.173628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:53:01.173628Z digest=sha256:b179874a2f6545871cc9b4a82b74d98cfe6e893e8a9b895a0426406d2a0e238b

Observation 37d98bee-ba20-498c-9bd2-637a8cb93ba7 · outbound

This paper cites Joint learning of attended zero-shot features and visual-semantic mapping,.

ResNetVLLM -- Multi-modal Vision LLM for the Video Understanding Task Joint learning of attended zero-shot features and visual-semantic mapping,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:53:01.853032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:53:01.177846Z digest=sha256:badc8f66c8be022560dc7435b9218f714d0173e80271f5a795fcdbce4ceb81fa

Observation 20177b18-fcaa-47fa-9e5d-44a34325ea9c · outbound

This paper cites A Review of Generalized Zero-Shot Learning Methods.

ResNetVLLM -- Multi-modal Vision LLM for the Video Understanding Task A Review of Generalized Zero-Shot Learning Methods

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-08-16T11:53:01.415927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:53:01.182322Z digest=sha256:8455ea774ec51d9fe5afc9e043987c1715cee1665acaeaf3cb50bfcc532248e4

Observation 89429e84-b8f8-4fef-98a3-303e0fb326eb · outbound

This paper cites Learning a deep embedding model for zero-shot learning,.

ResNetVLLM -- Multi-modal Vision LLM for the Video Understanding Task Learning a deep embedding model for zero-shot learning,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:53:01.838532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:53:01.186865Z digest=sha256:abeed8a65ec57ad456657f3d8ceb5753527956ad642d711fb8c3e7dfad6e5fb8

Observation 75935a25-512b-4e67-b893-7fbb54d1b0c1 · outbound

This paper cites Video question answering via gradually refined attention over appearance and motion,.

ResNetVLLM -- Multi-modal Vision LLM for the Video Understanding Task Video question answering via gradually refined attention over appearance and motion,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-16T11:53:01.191017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:53:01.191017Z digest=sha256:996be732088b9d2d9fe7e7914a82ab7972f6d7dd616cabdb3a8d3d6e45326502

Observation bb6fc5cb-8e23-4633-ba85-5cd65e039f23 · outbound

This paper cites Tgif-qa: Toward spatio- temporal reasoning in visual question answering,.

ResNetVLLM -- Multi-modal Vision LLM for the Video Understanding Task Tgif-qa: Toward spatio- temporal reasoning in visual question answering,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:53:01.814518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:53:01.195248Z digest=sha256:5c4d55acf0226fb69c09f4cb02b59b89f00b1adad8a3b45b7a1a67138fd41daf

Observation 7b1f0db8-c479-4411-be56-cef475623f71 · outbound

This paper cites Activitynet- qa: A dataset for understanding complex web videos via question answering,.

ResNetVLLM -- Multi-modal Vision LLM for the Video Understanding Task Activitynet- qa: A dataset for understanding complex web videos via question answering,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-16T11:53:01.199459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:53:01.199459Z digest=sha256:9d78463a1e055e64ec616dfff05ff8549016da4b40b21133d233315edeca7aa4

Observation 7656832f-b8f0-424e-924a-f9683f859924 · outbound

This paper cites Latent dirichlet allocation,.

ResNetVLLM -- Multi-modal Vision LLM for the Video Understanding Task Latent dirichlet allocation,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:53:01.791105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:53:01.203711Z digest=sha256:ffe9a117ab654717a9a0cde47e31a4bdc78e67e708afc054154bdfc016706718

Observation f78adb90-cac7-48a7-bdfb-27a7dd9b51ac · outbound

This paper cites Efficient estimation of word representations in vector space,.

ResNetVLLM -- Multi-modal Vision LLM for the Video Understanding Task Efficient estimation of word representations in vector space,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:53:01.777599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:53:01.208665Z digest=sha256:bb11e0a545fdbbfec07113b2227efab3a2d5540c479eaadb3221d9c92fbb6088

Observation 93315135-3f80-4e4e-b595-389b4aede8e1 · outbound

This paper cites Skip-thought vectors,.

ResNetVLLM -- Multi-modal Vision LLM for the Video Understanding Task Skip-thought vectors,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:53:01.764174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:53:01.213243Z digest=sha256:99533d3ecbca73445bc994a6149784b23ecadd1ee4c60d28688c1c9cf1b43c96

Observation 38ce0715-2bf0-48d1-bfe7-6d0469f846b7 · outbound

This paper cites Distributed representations of sentences and documents,.

ResNetVLLM -- Multi-modal Vision LLM for the Video Understanding Task Distributed representations of sentences and documents,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:53:01.750315Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:53:01.217700Z digest=sha256:a957811ee25062a44b4a27408c0c728f842cb1499ebdb13e872b9621cdabe986

Observation ff98a2ae-8787-494f-a374-ab359de1a45e · outbound

This paper cites A neural proba- bilistic language model,.

ResNetVLLM -- Multi-modal Vision LLM for the Video Understanding Task A neural proba- bilistic language model,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:53:01.736459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:53:01.222235Z digest=sha256:db37b71223ed37f4277a93b3916bc7e07f77713b160079ffe4ca982b27e2e131

Observation 0c37b4dd-802a-489e-a702-936c6219f618 · outbound

This paper cites Language Models are Few-Shot Learners.

ResNetVLLM -- Multi-modal Vision LLM for the Video Understanding Task Language Models are Few-Shot Learners

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-16T11:53:01.226446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:53:01.226446Z digest=sha256:95213a8f428b78d025b2cad1065c1a57cba65ff1a8a85816c7cbd3e14dde3ef3

Observation d05b9459-8a23-4a8a-bee2-7e4da5c94e42 · outbound

This paper cites Energy Tank-Based Policies for Robust Aerial Physical Interaction with Moving Objects.

ResNetVLLM -- Multi-modal Vision LLM for the Video Understanding Task Energy Tank-Based Policies for Robust Aerial Physical Interaction with Moving Objects

Reference 32

Resolution
metadata mismatch
local_arxiv, observed 2026-08-16T11:53:01.380764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:53:01.231020Z digest=sha256:36648d0e30a585d4ab273f5cd0d5ca15ae3261eb2036d9fe0181fcee461ca0b8

Observation d39ae4f1-3396-4ae7-b745-98df744e60bc · outbound

This paper cites OPT: Open Pre-trained Transformer Language Models.

ResNetVLLM -- Multi-modal Vision LLM for the Video Understanding Task OPT: Open Pre-trained Transformer Language Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-16T11:53:01.235484Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:53:01.235484Z digest=sha256:4f1d14f53932ae80fdb1e71038c36a9efe4ffb19b08132c942c858b0ba839103

Observation 70888efd-2d1a-42eb-b88f-26df1aec08d8 · outbound

This paper cites Flamingo: A visual language model for few-shot learning,.

ResNetVLLM -- Multi-modal Vision LLM for the Video Understanding Task Flamingo: A visual language model for few-shot learning,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:53:01.721945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:53:01.240051Z digest=sha256:942c58802c5656c48840412d84ef200cae8e96c1dea94ed0ae243a70a35b162b

Observation b99605de-16e8-49a2-8595-5a15dac97fdf · outbound

This paper cites Towards general purpose vision systems: An end-to-end task-agnostic vision-language architecture,.

ResNetVLLM -- Multi-modal Vision LLM for the Video Understanding Task Towards general purpose vision systems: An end-to-end task-agnostic vision-language architecture,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:53:01.708291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:53:01.244201Z digest=sha256:8b7db9694ed0b48db15e6ac6da40c555ae94e3fe635f3a24af097456f810a747

Observation 1318fc6c-a02a-4ad8-893f-e3974fc55ec3 · outbound

This paper cites Class-agnostic object detection with multi-modal transformer,.

ResNetVLLM -- Multi-modal Vision LLM for the Video Understanding Task Class-agnostic object detection with multi-modal transformer,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:53:01.694600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:53:01.248550Z digest=sha256:6c2bf0a15770ad5db27c8554e99e025314721cd34f2aaaaf9abb5d025d0c5de9

Observation ad11aa57-caf0-4881-8015-e5d4f645d8a4 · outbound

This paper cites Bridg- ing the gap between object and image-level representations for open- vocabulary detection,.

ResNetVLLM -- Multi-modal Vision LLM for the Video Understanding Task Bridg- ing the gap between object and image-level representations for open- vocabulary detection,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:53:01.679889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:53:01.253078Z digest=sha256:26dc3de9ec9362e5c15a9b5fc38a511798a4dd8b59c64f9d1184e9951a00e4d0

Observation c466c31f-8346-4db8-8666-49e04d9a59b0 · outbound

This paper cites Open-vocabulary semantic segmentation with mask-adapted clip,.

ResNetVLLM -- Multi-modal Vision LLM for the Video Understanding Task Open-vocabulary semantic segmentation with mask-adapted clip,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:53:01.666490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:53:01.257286Z digest=sha256:b6290279416d3767e1418a990830cda66f09a03aa788aece672cc02d005446d0

Observation 83e32dad-0985-4a08-afed-7c6f3a8c4d07 · outbound

This paper cites Language-grounded indoor 3d semantic segmentation in the wild,.

ResNetVLLM -- Multi-modal Vision LLM for the Video Understanding Task Language-grounded indoor 3d semantic segmentation in the wild,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:53:01.652301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:53:01.261324Z digest=sha256:c59e9f0493ef2ce417c1b0f55486ff120c40861c0307ee850a90b42e3d529ebd

Observation bdd060f2-07d9-4739-9535-d14d05633eaf · outbound

This paper cites Expanding language-image pretrained models for general video recognition,.

ResNetVLLM -- Multi-modal Vision LLM for the Video Understanding Task Expanding language-image pretrained models for general video recognition,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:53:01.637665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:53:01.265658Z digest=sha256:ed72b29c185aa3751cd08a359426521f6aacc25fe790ea6178f6e859779c9708

Observation ec746aca-0426-400e-9372-d412337a601a · outbound

This paper cites ActionCLIP: A New Paradigm for Video Action Recognition.

ResNetVLLM -- Multi-modal Vision LLM for the Video Understanding Task ActionCLIP: A New Paradigm for Video Action Recognition

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-16T11:53:01.269990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:53:01.269990Z digest=sha256:cfc1db200f29625610b4729ddd175660144ef4f24100b47f3d6e0cd4b6a59c0a

Observation a46ef225-0370-40f3-8308-8689c32b7d34 · outbound

This paper cites Finetuned clip models are efficient video learners,.

ResNetVLLM -- Multi-modal Vision LLM for the Video Understanding Task Finetuned clip models are efficient video learners,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:53:01.622589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:53:01.274197Z digest=sha256:4cf91454de65918f1b5c6d5b3f1dbb81b56db34ba5c135193623bb87023fe0be

Observation 8417e37e-6722-40b1-9afb-9a3f9743a26e · outbound

This paper cites Deep residual learning for image recognition,.

ResNetVLLM -- Multi-modal Vision LLM for the Video Understanding Task Deep residual learning for image recognition,

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-16T11:53:01.278352Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:53:01.278352Z digest=sha256:f13d720b7eec23a18510ff5aa9b9e2d907a8b4698721329d3126fbfdfd9d5931

Observation e0eebc19-9393-4a68-9cad-bba3fedf1478 · outbound

This paper cites Temporal segment networks: Towards good practices for deep action recognition,.

ResNetVLLM -- Multi-modal Vision LLM for the Video Understanding Task Temporal segment networks: Towards good practices for deep action recognition,

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-16T11:53:01.282427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:53:01.282427Z digest=sha256:99ba7980a4139719bd557bed00f6f9e477aba79d665f7eb6e8e8815ef35b2c85

Observation 0604e75d-9422-4794-ab3b-cc9e8cf058a7 · outbound

This paper cites Activi- tynet: A large-scale video benchmark for human activity understanding,.

ResNetVLLM -- Multi-modal Vision LLM for the Video Understanding Task Activi- tynet: A large-scale video benchmark for human activity understanding,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:53:01.590115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:53:01.287245Z digest=sha256:970f7a698689c953b5465f9847d3f3d5bbda629f70c47dfeac008371d844502a

Observation 0ba76b65-e53d-49a9-8a97-f5deaed3158c · outbound

This paper cites BT-Adapter: Video Conversation is Feasible Without Video Instruction Tuning.

ResNetVLLM -- Multi-modal Vision LLM for the Video Understanding Task BT-Adapter: Video Conversation is Feasible Without Video Instruction Tuning

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-16T11:53:01.291124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:53:01.291124Z digest=sha256:fd06d7dcca22debe91c90759fdbb9b4e3ff17494829e08a19a3d5d06ccc5c22e

Observation c0097f47-a2fb-420e-999e-701f5f012d1a · outbound

This paper cites Zero-shot video question answering via frozen bidirectional language models,.

ResNetVLLM -- Multi-modal Vision LLM for the Video Understanding Task Zero-shot video question answering via frozen bidirectional language models,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:53:01.575169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:53:01.295207Z digest=sha256:dde6cd589b3b91f203b9d37ac785dd9c2e7bc03cefbb4075e441c3e4ab5100f6

Pith citing papers

No inbound Pith citation observations are available.