Pith. sign in

Paper Citation Record · LEDGER

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding

As of 14 August 2026, this Paper Citation Record lists 80 of 80 outbound references and 0 inbound Pith citation observations for arXiv:2411.12355.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.12355 v2

Coverage vector

measured 80 of 80 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T17:43:03.400998Z

measured 80 of 80 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

80 of 80 outbound references displayed

  • verified exact0
  • verified fuzzy44
  • unresolved35
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2192909b-a3ec-40ac-aa7f-5a5b0ac5a6cd · outbound

This paper cites an unresolved cited work.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-12T17:43:04.840167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:43:02.999634Z digest=sha256:c584af6d8da36306f7da2c686565a409022d2aef4ea0b19de163f7188008a07a

Observation ccf63729-4ef7-4c31-9099-15657401fcf3 · outbound

This paper cites Textvqa: Towards understanding of visible and invisible text in images.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Textvqa: Towards understanding of visible and invisible text in images

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:04.821791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:43:03.005333Z digest=sha256:46dd15f38f7f74364a73f831c9b6ef52cc08f1cccd0ea42f28b96ee57e10a0ef

Observation a7df40f1-6a21-465e-8c84-95e39da5afaa · outbound

This paper cites MiniGPT4-Video: Advancing Multimodal LLMs for Video Understanding with Interleaved Visual-Textual Tokens.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding MiniGPT4-Video: Advancing Multimodal LLMs for Video Understanding with Interleaved Visual-Textual Tokens

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T17:43:03.010674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:43:03.010674Z digest=sha256:35790e67ccbd75ef763b435f96965d074406588a8906c7a2eed5c77a2dc5f39c

Observation ae255e3c-586c-4dc1-aa72-d127c3ac2d32 · outbound

This paper cites Frozen in time: A joint video and image encoder for end-to- end retrieval.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Frozen in time: A joint video and image encoder for end-to- end retrieval

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:04.802082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:43:03.016186Z digest=sha256:30425ee85ff9d924904818621af9e8ff57acde90593cb4728fd927cff0d55df5

Observation 036feb4b-a3b7-4b69-93c4-d10d7eba5f3c · outbound

This paper cites Frozen in time: A joint video and image encoder for end-to- end retrieval.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Frozen in time: A joint video and image encoder for end-to- end retrieval

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:04.782344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:43:03.021259Z digest=sha256:b8156deced4842111504c19a6228d165cbe86f22f1edba0ff68d1e6c5f429f8b

Observation 47c3774e-a343-4e4e-9a7a-3209d40cd88d · outbound

This paper cites Learning with Differentiable Perturbed Optimizers.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Learning with Differentiable Perturbed Optimizers

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T17:43:03.026435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:43:03.026435Z digest=sha256:e523d2bfdf0807478335652317884b21e91eb7f42104cf6fe7e7076c5ea36942

Observation 169e7c30-2c2d-47b9-b16c-25b842bf533a · outbound

This paper cites Activitynet: A large-scale video bench- mark for human activity understanding.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Activitynet: A large-scale video bench- mark for human activity understanding

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:04.762216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:43:03.032611Z digest=sha256:a4e2e6e17eba6ddedd421cd419bd06161352e91aef17e6dc2c07accd460cde01

Observation e2174da7-e12b-41df-aeb8-06e665057820 · outbound

This paper cites Chen and William B.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Chen and William B

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:04.741759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:43:03.037293Z digest=sha256:93a511f80565b38a14e898a16a6216bc64db2500e4abccdb2e3702106333bd2a

Observation 8c0bab9f-0d08-4aef-b355-b32aadb75f55 · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T17:43:03.047959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:43:03.047959Z digest=sha256:d6df0251df3df05e22afa31c9faad76b0ec2d32fa1457df4c69f549eb30eed20

Observation 7c0a61f2-f1a3-46bd-9800-de413e0a100f · outbound

This paper cites Gonzalez, Ion Stoica, and Eric P.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Gonzalez, Ion Stoica, and Eric P

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T17:43:03.053816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:43:03.053816Z digest=sha256:69f895a07d2727540f659557d37537dd0c19121925dc4bd17f375ee889769d50

Observation 53588708-e7e8-44db-86f2-a24f74c2d2c5 · outbound

This paper cites Palm: Scaling language modeling with pathways.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Palm: Scaling language modeling with pathways

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T17:43:03.058756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:43:03.058756Z digest=sha256:8392f2d5f7af83f404a04713f3cde0004ecfe8d1b5165617e09edae6ddedecb0

Observation a94e7e06-e704-43d2-8ba4-b05789559123 · outbound

This paper cites an unresolved cited work.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-12T17:43:04.677150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:43:03.063662Z digest=sha256:4e7afb16017c56758be053728f08316530fd19209b7e0b6f85d654e4b43757d8

Observation 8923c3c4-b5a4-4d60-b4ec-53ed7bae0b2e · outbound

This paper cites EgoQA: Egocentric question an- swering.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding EgoQA: Egocentric question an- swering

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:04.659167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:43:03.068378Z digest=sha256:67b505a0e90acc0cc8078252a793494e93704cb70bf91ec5b7444bb52255cf84

Observation 5178ef37-024c-4463-8b7d-feb5f9a4adf1 · outbound

This paper cites Study on density peaks clustering based on k-nearest neighbors and principal component analysis.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Study on density peaks clustering based on k-nearest neighbors and principal component analysis

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:04.640474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:43:03.073060Z digest=sha256:ecf42f2384657e3e6756247229c7bb465abd2e90bfad9c591be9d180c91216ff

Observation 114aef3b-3e26-4775-92a2-779f5607bec3 · outbound

This paper cites EV A: exploring the limits of masked visual representation learning at scale.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding EV A: exploring the limits of masked visual representation learning at scale

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:04.622796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:43:03.077586Z digest=sha256:d0238f9b08a7a93c131654250f0438b012a4f3f74d2ac95c41f1de7b98b2720e

Observation 7a073534-89e0-4cd5-8319-8f235bfaaaef · outbound

This paper cites Making the V in VQA matter: Elevating the role of image understanding in visual question answering.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Making the V in VQA matter: Elevating the role of image understanding in visual question answering

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:04.604589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:43:03.082185Z digest=sha256:3666804ddb2eb57f107f3de7bc514ecf7dae3e76ef8c74c7abae57f7bc6f2077

Observation 2b2efb6a-fb1c-44e9-ad25-53a1c8e3cfd0 · outbound

This paper cites Semantic-aware modular capsule routing for visual question answering.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Semantic-aware modular capsule routing for visual question answering

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:04.587504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:43:03.086860Z digest=sha256:a1fc893e1dd0fc2ba2343e04773f9671900020cfea514a9207db4fcb1a24414f

Observation 967e23a5-a606-4c25-8945-23b63da30962 · outbound

This paper cites MA-LMM: memory-augmented large multimodal model for long-term video understanding.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding MA-LMM: memory-augmented large multimodal model for long-term video understanding

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:04.568683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:43:03.091661Z digest=sha256:6b69b166c7e8671489bf5d4c56ab854ae69fb2e399f8a62c0d53bb16bda6f266

Observation 5de07298-2e8b-4d31-94ad-f6ed51167c0a · outbound

This paper cites Ma-lmm: Memory-augmented large multimodal model for long-term video understanding.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Ma-lmm: Memory-augmented large multimodal model for long-term video understanding

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:04.551934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:43:03.096166Z digest=sha256:56a5103eb24345d90fb9461bc22784d6013a3fd7c2a5cf5683daa1059065ad93

Observation 6d4b9996-60c8-438c-b156-f73174c9c75e · outbound

This paper cites Activitynet: A large-scale video bench- mark for human activity understanding.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Activitynet: A large-scale video bench- mark for human activity understanding

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:04.535708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:43:03.100797Z digest=sha256:b66fa9291aca9ff6919fb97fbf69724fc36b88b7f24c8e1442a7594081d4bb92

Observation 3d217d76-55f0-4919-988e-a7158c81f2f2 · outbound

This paper cites an unresolved cited work.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-12T17:43:04.518297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:43:03.105442Z digest=sha256:5f9ed6939ba078ded865ab5fef7e706fb794554f9f5feeed2a4720ab39429107

Observation 02a8aba9-ed37-42e7-a542-a2bc61b21d20 · outbound

This paper cites VTimeLLM: Empower LLM to Grasp Video Moments.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding VTimeLLM: Empower LLM to Grasp Video Moments

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T17:43:03.110131Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:43:03.110131Z digest=sha256:accc853b6d3817e4785b525e91ecdcee73c86ecb587c6af4631765e30b41b16c

Observation a4a3e019-fd4d-42d8-8cec-21765358ad1e · outbound

This paper cites Ng, Hongqiang Rong, and Zichen Li.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Ng, Hongqiang Rong, and Zichen Li

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:04.501170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:43:03.115496Z digest=sha256:e93b19db69c3fda074c63eb302ff260b77ad48eb6a384243c768b92970d8df9a

Observation 1c862dea-09fb-40c2-8280-c60f3386f8ad · outbound

This paper cites Gqa: A new dataset for real-world visual reasoning and compositional questions.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Gqa: A new dataset for real-world visual reasoning and compositional questions

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:04.484544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:43:03.120120Z digest=sha256:54e036767cfa1a7f0e74f13b4a47948ad12b4c0bbbbc06032009f490c42f57ca

Observation a63abfb8-567c-49a0-aabe-73c348fdd3a4 · outbound

This paper cites Chat-UniVi: Unified Visual Representation Empowers Large Language Models with Image and Video Understanding.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Chat-UniVi: Unified Visual Representation Empowers Large Language Models with Image and Video Understanding

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T17:43:03.124790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:43:03.124790Z digest=sha256:fb2b8ba203cf79b756c33f7892419ec0dcfae3e304688c7ffe277cdf5cfecd1f

Observation 2611f948-041c-4fcb-8df2-73f54e44b3ed · outbound

This paper cites Scaling Laws for Neural Language Models.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Scaling Laws for Neural Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T17:43:03.130132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:43:03.130132Z digest=sha256:694c390befbab345d3b6e8460293a9a289f6509adfcad010fd184946cdcf6970

Observation 7b32a324-e24f-4f7b-ac0d-ebe87d4d61ba · outbound

This paper cites an unresolved cited work.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-12T17:43:04.463976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:43:03.135145Z digest=sha256:3129b8b7451d8d877a328f167ce27cb4622c87b666e27de220f44e1cce136d03

Observation d1aedadd-e8e9-4b7c-bc4b-b87f8f061e44 · outbound

This paper cites Kingma and Jimmy Ba.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Kingma and Jimmy Ba

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:04.444757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:43:03.140583Z digest=sha256:6a80a1f68a319ec49d847e1a6e0ce9b8b00ede1c61a0c80f44b05b5ac1fed6e3

Observation 1358aabb-24c5-4dfd-84b4-abf6534493aa · outbound

This paper cites Visual genome: Connecting language and vision using crowdsourced dense image annotations.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Visual genome: Connecting language and vision using crowdsourced dense image annotations

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:04.428366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:43:03.145305Z digest=sha256:67bdbff306bd8d469b52dd316d0197e5a0fab4135271b9f69db2b15f5ae86f20

Observation fe1aeb9e-49b2-4b88-8ada-5996889d545e · outbound

This paper cites Visual genome: Connecting language and vision using crowdsourced dense image annotations.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Visual genome: Connecting language and vision using crowdsourced dense image annotations

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:04.411158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:43:03.150322Z digest=sha256:a9ffdcf62939535bdd4f1dfd321778b71c8f2ab8ed1273f6d703db92fb6e5b3f

Observation 8dd26375-64c0-4b7c-87be-1d82ecc6d9e3 · outbound

This paper cites VideoChat: Chat-Centric Video Understanding.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding VideoChat: Chat-Centric Video Understanding

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T17:43:03.155428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:43:03.155428Z digest=sha256:4ba722677f63128051d1cda7d208010bc8a41f1c819a0d11bbd3a18e178f8d60

Observation 4b36e757-7f2d-47c3-b746-91d1986490f7 · outbound

This paper cites MVBench: A Comprehensive Multi-modal Video Understanding Benchmark.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding MVBench: A Comprehensive Multi-modal Video Understanding Benchmark

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T17:43:03.160467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:43:03.160467Z digest=sha256:6d5901268ca9dea6a357094e8684eaef88bd6fbe5a4ba401b8b46e436d3a2330

Observation d816dd29-779c-4244-9cf9-18ca9a8829ea · outbound

This paper cites Scienceqa: A new dataset for science question answering.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Scienceqa: A new dataset for science question answering

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:04.392817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:43:03.165529Z digest=sha256:c65f3c02aaeb7e1d6127510dd2388b7530e09097ed7edeeea5e7745f6c5bd044

Observation afc0463e-dede-48a9-9cd3-afd740fe1de1 · outbound

This paper cites Learning dynamic routing for semantic segmentation.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Learning dynamic routing for semantic segmentation

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:04.373424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:43:03.170423Z digest=sha256:d2983bc44caa2271193b5acd2afa9f008c49404ae7e8c70f48d3ab6427c76268

Observation eb7e9455-6eaf-461f-8f19-563a7e6d1ab0 · outbound

This paper cites LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-12T17:43:03.175123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:43:03.175123Z digest=sha256:5be9407978f68b9eedc6283aa0a30ef991831229492ece7726323e9440e8f459

Observation 66d3c174-85f9-447b-a60a-1f540dc777cb · outbound

This paper cites Video-llava: Learning united visual representa- tion by alignment before projection.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Video-llava: Learning united visual representa- tion by alignment before projection

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:04.351160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:43:03.180928Z digest=sha256:4d80db797c677e6f8c5a50779d598a1a55265b4fd91b670b56b06300526dec56

Observation c1e97409-d45b-42bf-98a6-7a84d74224b8 · outbound

This paper cites Video-LLaVA: Learning United Visual Representation by Alignment Before Projection.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-12T17:43:03.186135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:43:03.186135Z digest=sha256:ca99a36afece154d629c93f2d8c98661d4dd079e651bede0ebb5b4a00b7459f5

Observation 90fcb0c1-438f-463f-8bb4-6fd700a3c9d0 · outbound

This paper cites Vila: On pre-training for visual language models.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Vila: On pre-training for visual language models

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:04.333933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:43:03.191688Z digest=sha256:ed186c1a5ab1ec764e261a86104b7074bc3b57c6c3c034337e0556fc7c031805

Observation 79bff67f-cb71-4e27-9c1b-2d75451d3a95 · outbound

This paper cites Lawrence Zitnick.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Lawrence Zitnick

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:04.315970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:43:03.196680Z digest=sha256:4c453b68559f000268249bbb90a6278f58ef520a3c53b7a950c6b72b6d670911

Observation 25d71ced-da19-4f2f-a516-afa49035b62f · outbound

This paper cites Improved baselines with visual instruction tuning, 2023.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Improved baselines with visual instruction tuning, 2023

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-12T17:43:03.203557Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:43:03.203557Z digest=sha256:e127bd4b2c8c99436119d9c481691380c56183285005f71df6e7f90dc4dba98e

Observation de0d25fb-e774-4205-9732-0b5a640fbf72 · outbound

This paper cites World Model on Million-Length Video And Language With Blockwise RingAttention.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding World Model on Million-Length Video And Language With Blockwise RingAttention

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-12T17:43:03.208853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:43:03.208853Z digest=sha256:8439a83d7fb50f1fc5dbbab3c2c798665f9ce8217a83839e58dfc7cef0c1892f

Observation 60a8e0cb-1022-4047-a7ed-825647eab5e4 · outbound

This paper cites BT-Adapter: Video Conversation is Feasible Without Video Instruction Tuning.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding BT-Adapter: Video Conversation is Feasible Without Video Instruction Tuning

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-12T17:43:03.214641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:43:03.214641Z digest=sha256:ea3f4b83249a3b98a42caa7df10b1340626e73e277414932b46fb2c8dc29ad2f

Observation 1593324f-320e-4af4-9cb9-cfc8b1ed7bc7 · outbound

This paper cites ST-LLM: Large Language Models Are Effective Temporal Learners.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding ST-LLM: Large Language Models Are Effective Temporal Learners

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-12T17:43:03.220788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:43:03.220788Z digest=sha256:f655664d829bba7aa2953fcec2ffa0de5e556f9b5ba4a84b0a438f02b2bbf47d

Observation 2af014f9-11db-4cdd-aa50-d2a8f3d2c860 · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Learn to explain: Multimodal reasoning via thought chains for science question answering

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-12T17:43:03.226223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:43:03.226223Z digest=sha256:866f0bfb4b0edf2a72cf83a2a2ea8b91da8b78d402acc7c9e36f9680c7d5ea7e

Observation 2ffbb597-19fb-4961-93fd-8e4f31a16a59 · outbound

This paper cites Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-12T17:43:03.231069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:43:03.231069Z digest=sha256:dcfcc2534e97a96007694f9db00361a81cc9cee4a4e2be8f8b4e50b30ab44c1c

Observation b8a8055a-73b3-44f0-91dd-8e70d97aa0ee · outbound

This paper cites Some methods for classification and anal- ysis of multivariate observations.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Some methods for classification and anal- ysis of multivariate observations

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:04.272680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:43:03.236284Z digest=sha256:45a9d255f1c45980af058b2a11d166d7c619054106b2f05be1ba507c2d5b7c37

Observation 1f1405bd-bdfa-4a08-b6e1-179d714a455f · outbound

This paper cites Yuille, and Kevin Murphy.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Yuille, and Kevin Murphy

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:04.253109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:43:03.241058Z digest=sha256:7fca16d823f3b474222d7acd0860ce92e7e3ff2a59bae79b6f091c397c80046c

Observation fbd14a40-e224-4093-8831-e1162e9cf647 · outbound

This paper cites Ocr-vqa: Visual question answering by reading text in images.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Ocr-vqa: Visual question answering by reading text in images

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-12T17:43:03.245850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:43:03.245850Z digest=sha256:a4813aeb58ac66d5a06a630e6af344b1924c474114f1137c86a0b3229cdc53c2

Observation ba16ab19-d20a-4e06-9572-00797ee7c6ab · outbound

This paper cites Webvidqa: A large-scale dataset for video question answering.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Webvidqa: A large-scale dataset for video question answering

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:04.224590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:43:03.251180Z digest=sha256:d9b3b81a9f73058490913a2c7efb34aaf31630f5e5be59b0eb3d0523387798c9

Observation 6ea3099d-e5d3-48e4-be36-b99bf90505b1 · outbound

This paper cites Introducing chatgpt.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Introducing chatgpt

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:04.208955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:43:03.256692Z digest=sha256:12db2ba1ab4afe593df11c3b3428c1499104661927a8cc4f5a3b24e64a8b3299

Observation b9e6b080-9048-43f3-8b18-406637cedd42 · outbound

This paper cites GPT-4o system card, 2024.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding GPT-4o system card, 2024

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:04.192873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:43:03.261844Z digest=sha256:08af09e745bd0abef2dbfadb562229ac47b77b6228f999627f6abd6411a3fb29

Observation 6ecae718-bdbf-4511-a75c-0ff1b842447e · outbound

This paper cites Learning transferable visual models from natural language supervision.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Learning transferable visual models from natural language supervision

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:04.177914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:43:03.267235Z digest=sha256:547b79861536750e304f3ecf19dbf5ee17114266af4c927d87f52b7d7e2b5ca8

Observation b2a3c4cd-944a-46b1-a257-bab491b451bf · outbound

This paper cites Improving language understanding by gener- ative pre-training.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Improving language understanding by gener- ative pre-training

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:04.162143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:43:03.272111Z digest=sha256:76580c5b6a51f1827a344d4d930a9c52c55e5c27dfd82ab31b273dfd031b6325

Observation 80ab209e-d25d-468f-b4f5-f2c99e816a88 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:04.146210Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:43:03.277089Z digest=sha256:d68146453be93d608e3bd206939fde005611025262c444ec39481150052fe497

Observation 849ce177-89ab-4c57-84b2-263947569873 · outbound

This paper cites Timechat: A time-sensitive multimodal large language model for long video understanding.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Timechat: A time-sensitive multimodal large language model for long video understanding

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:04.129764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:43:03.281962Z digest=sha256:90411700b0b95dd3beb8ae5cef0a6b068e7857fd8bd489af1e53b531c5b5ec49

Observation c8db052f-139d-4721-92db-375d350333a0 · outbound

This paper cites Massof Sarah L.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Massof Sarah L

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:04.112583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:43:03.288087Z digest=sha256:a93720c1fa764c5a231ab50d7c1d792bb969c49fba068e4bf14973b13ca8890d

Observation e74eb7fb-e81a-4a1a-aedb-f5c5f9d1a06d · outbound

This paper cites A-okvqa: A bench- mark for visual question answering using world knowledge.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding A-okvqa: A bench- mark for visual question answering using world knowledge

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:04.095257Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:43:03.292697Z digest=sha256:9c1570e4133de008d3670ecdf254d2329c3de5efe533a6168ad5899b60927f59

Observation 35c00115-a116-4de1-9640-eabf185e24a1 · outbound

This paper cites Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:04.078763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:43:03.297505Z digest=sha256:91c5c6bc7dd07dd54246524f758fdd8d09aa413c170ab0b26fc4ccf33bc657a4

Observation 04cbea46-a537-4ccb-a7c1-a20ce0943b2b · outbound

This paper cites Textcaps: a dataset for image captioning with reading comprehension.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Textcaps: a dataset for image captioning with reading comprehension

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:04.059597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:43:03.302047Z digest=sha256:67a574d30cfcaa240dd4513dc3a4811cd1a2a413684c9ec74a8c5e78e287d94d

Observation 07fcd0e9-4a9d-48f4-913e-16a82ba77ff9 · outbound

This paper cites MovieChat: From Dense Token to Sparse Memory for Long Video Understanding.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding MovieChat: From Dense Token to Sparse Memory for Long Video Understanding

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-12T17:43:03.306611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:43:03.306611Z digest=sha256:46906cd58c5f68af9cab14d078dd9a7263e7476438af1c3dcc8e63b926be4b38

Observation 244acf83-61ea-4a6b-9dd9-53248aa8ebca · outbound

This paper cites Moviechat: From dense token to sparse memory for long video understanding.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Moviechat: From dense token to sparse memory for long video understanding

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:04.038794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:43:03.311747Z digest=sha256:49a4b0e7e2e249c9211a3cdfb89d7770c2f46fdd8d89d99413654022b4d03554

Observation adc1f000-41c0-4e15-beea-8be646f352d0 · outbound

This paper cites DreamFrame: Enhancing Video Understanding via Automatically Generated QA and Style-Consistent Keyframes.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding DreamFrame: Enhancing Video Understanding via Automatically Generated QA and Style-Consistent Keyframes

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-12T17:43:03.316445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:43:03.316445Z digest=sha256:bb09e4d4856986d8624cd7cfcb536184da3989ac3e151d0fd0b919658fe98d1f

Observation f27cfd0f-7d32-4381-9795-91a3214f24e6 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding LLaMA: Open and Efficient Foundation Language Models

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-12T17:43:03.321491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:43:03.321491Z digest=sha256:78c9781b21cddc5ca12d2b9ed281ba2fc8d528d818375b4ea7472c924a70cfd6

Observation 43a75b67-c0a8-435b-9aac-d5c43325d460 · outbound

This paper cites Ocrvqa: A new dataset for optical character recognition in visual question answering.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Ocrvqa: A new dataset for optical character recognition in visual question answering

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:04.020499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:43:03.326429Z digest=sha256:04a680037344a3d28e30334fdcfb11293aef365f4809e6efbd1e05cecd6466cb

Observation 7e892c22-933f-4318-965b-dc19ed6cf8a9 · outbound

This paper cites LVBench: An Extreme Long Video Understanding Benchmark.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding LVBench: An Extreme Long Video Understanding Benchmark

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-12T17:43:03.331072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:43:03.331072Z digest=sha256:7b28e36c6e4362315ba133e6dff54c721fc876fca01420870e8bcabb0328b7be

Observation 97f740c3-73c6-4c9e-8958-67f8bdfdf36e · outbound

This paper cites VideoHallucer: Evaluating Intrinsic and Extrinsic Hallucinations in Large Video-Language Models.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding VideoHallucer: Evaluating Intrinsic and Extrinsic Hallucinations in Large Video-Language Models

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-12T17:43:03.336459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:43:03.336459Z digest=sha256:3ea1f6e2a22c7274e3c81443491e3f2fd43c1123fe2691b6c42bdc0e684506b2

Observation bd29c8d1-fd60-4a56-b0ad-e2dc3962211b · outbound

This paper cites Videollamb: Long video understanding with recurrent mem- ory bridges.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Videollamb: Long video understanding with recurrent mem- ory bridges

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:04.003342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:43:03.341593Z digest=sha256:ffb596423bd29c6ae7278da833afda8457de3d6d897bcd636ed0f11e52ee469f

Observation cc9702e0-bbdb-40bd-a4b7-33a4e5b83220 · outbound

This paper cites FreeVA: Offline MLLM as Training-Free Video Assistant.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding FreeVA: Offline MLLM as Training-Free Video Assistant

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-12T17:43:03.346619Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:43:03.346619Z digest=sha256:326948271115a114ce41101eb392106cb13a58172f46e144716ffe47403b5129

Observation 0d789eab-03a8-4918-a2c7-8a4e8d87ea3a · outbound

This paper cites Davis, Kristen Grauman, and Rog´erio Schmidt Feris.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Davis, Kristen Grauman, and Rog´erio Schmidt Feris

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:03.986333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:43:03.352245Z digest=sha256:0a9983cb210be29a2db0b66c8d081c56b2d7cd4ba5382bfadc2472574cc72587

Observation c824db95-0761-4a49-8aef-c606a9170226 · outbound

This paper cites Deep learning for video classification and captioning.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Deep learning for video classification and captioning

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:03.969334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:43:03.356916Z digest=sha256:7c230f615c7f6bce1d73ee568ed5744df570a102e1b433fa6a6b5e78269a7ea6

Observation b0b61000-7204-46e3-9c89-70cbe5eb90f7 · outbound

This paper cites Msr-vtt: A large video description dataset for bridging video and language.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Msr-vtt: A large video description dataset for bridging video and language

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-12T17:43:03.361819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:43:03.361819Z digest=sha256:11936d3317eea5173e3c68a4a5a462e32b3d92b85a31afdb1022b38355776ff2

Observation b417a12e-c613-4623-b461-5dcfb2fc48c0 · outbound

This paper cites MSR-VTT: A large video description dataset for bridging video and language.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding MSR-VTT: A large video description dataset for bridging video and language

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:03.940819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:43:03.366481Z digest=sha256:ae6d123cf52f7b19382b8fd08d1b619733de6d6a5c22384cf5f5cdbdd1811926

Observation 580d49b8-86fd-4237-b6e6-0aaf0f50a8b1 · outbound

This paper cites PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-12T17:43:03.371221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:43:03.371221Z digest=sha256:1ca7981839612cb642bc7232c93d71d0538f9347d38d1258885418b1b329dd5b

Observation 0886b48f-09a5-4d1f-a63a-4601f2da034d · outbound

This paper cites CLEVRER: CoLlision Events for Video REpresentation and Reasoning.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding CLEVRER: CoLlision Events for Video REpresentation and Reasoning

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-12T17:43:03.376477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:43:03.376477Z digest=sha256:a300728dfb38ed500125c2b76c90257fe2891da80792e2bedd6b8121d89038d8

Observation 86ed1b74-7c25-43e6-9888-5a20b69ec836 · outbound

This paper cites Tenenbaum.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Tenenbaum

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:03.924319Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:43:03.381552Z digest=sha256:f503f96a1c6c8a43ca04dd705d716a8192dbf56f9016018f55677ca85cd79a2f

Observation ac5110db-7965-434d-a3c8-be105974b135 · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-12T17:43:03.386720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:43:03.386720Z digest=sha256:0db88c8f87d2dd014ca43ad049756938219714ddc648992a8280d253bcd8765d

Observation 5cbfdf1e-ed6a-4579-8036-f4c7b9d9ea87 · outbound

This paper cites Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-12T17:43:03.391576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:43:03.391576Z digest=sha256:d5bbb7d1624e84c0181a3c7b4354d0c11f40077184e8f3a9512371e6add92c03

Observation f67af649-7e95-4e51-9057-9c8392cd8a03 · outbound

This paper cites Llama- adapter: Efficient fine-tuning of language models with zero- init attention.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Llama- adapter: Efficient fine-tuning of language models with zero- init attention

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:03.906623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:43:03.396276Z digest=sha256:618428ede7a2a7a9e187a31f43818d7d7c7be8d3bff0aa469699e413c5769321

Observation f7d1270e-acba-4a7c-b2a9-061431d94cc9 · outbound

This paper cites Please Carefully Think.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Please Carefully Think

Reference 79

Resolution
malformed identifier
raw_fallback, observed 2026-08-12T17:43:03.887845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:43:03.400998Z digest=sha256:0985624d88d963a15472063f16be190c5ce2ad83ee9f7d6a68f7f2b7fb12549e

Observation 831ab334-8259-405f-a2cb-5c78bfdc1331 · outbound

This paper cites an unresolved cited work.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Unresolved cited work

Reference 200

Resolution
unresolved
raw_fallback, observed 2026-08-12T17:43:04.719668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:43:03.042499Z digest=sha256:8c35584fff37c7d4ee16304ed7549d5ba856e7102bfb435a2aaaf55a9dabdd5a

Pith citing papers

No inbound Pith citation observations are available.