Pith. sign in

Paper Citation Record · LEDGER

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding

As of 14 August 2026, this Paper Citation Record lists 80 of 80 outbound references and 0 inbound Pith citation observations for arXiv:2411.12355.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.12355 v2

Coverage vector

measured 80 of 80 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T17:43:03.400998Z

measured 80 of 80 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

80 of 80 outbound references displayed

  • verified exact0
  • verified fuzzy44
  • unresolved35
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2192909b-a3ec-40ac-aa7f-5a5b0ac5a6cd · outbound

This paper cites an unresolved cited work.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-12T17:43:04.840167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:43:02.999634Z digest=sha256:4648d366a53d52e7a920637bcdb1f2cb7e9d02e02927a20c030030755754d551

Observation ccf63729-4ef7-4c31-9099-15657401fcf3 · outbound

This paper cites Textvqa: Towards understanding of visible and invisible text in images.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Textvqa: Towards understanding of visible and invisible text in images

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:04.821791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:43:03.005333Z digest=sha256:e77b375f2b00cfeaace429e1eef780aa81d7d93fff74721c73fade372c74c91c

Observation a7df40f1-6a21-465e-8c84-95e39da5afaa · outbound

This paper cites MiniGPT4-Video: Advancing Multimodal LLMs for Video Understanding with Interleaved Visual-Textual Tokens.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding MiniGPT4-Video: Advancing Multimodal LLMs for Video Understanding with Interleaved Visual-Textual Tokens

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T17:43:03.010674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:43:03.010674Z digest=sha256:f0f9dc703c8f2241b22c3c0a85c2708ef34408d86035c736f05247a3b1246f70

Observation ae255e3c-586c-4dc1-aa72-d127c3ac2d32 · outbound

This paper cites Frozen in time: A joint video and image encoder for end-to- end retrieval.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Frozen in time: A joint video and image encoder for end-to- end retrieval

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:04.802082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:43:03.016186Z digest=sha256:9ad51acd8271bbd9da4e067d9d38bc02389252804bcf0e31e15b3851c3cbd3c5

Observation 036feb4b-a3b7-4b69-93c4-d10d7eba5f3c · outbound

This paper cites Frozen in time: A joint video and image encoder for end-to- end retrieval.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Frozen in time: A joint video and image encoder for end-to- end retrieval

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:04.782344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:43:03.021259Z digest=sha256:e3ea373804777fd7f13a85de370500a53d8b24ba99ddac368228d5c5667c3e45

Observation 47c3774e-a343-4e4e-9a7a-3209d40cd88d · outbound

This paper cites Learning with Differentiable Perturbed Optimizers.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Learning with Differentiable Perturbed Optimizers

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T17:43:03.026435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:43:03.026435Z digest=sha256:40448882e330137520e8ed9e93c91a245874ff8bf705fe0bb270d3cc13c554d0

Observation 169e7c30-2c2d-47b9-b16c-25b842bf533a · outbound

This paper cites Activitynet: A large-scale video bench- mark for human activity understanding.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Activitynet: A large-scale video bench- mark for human activity understanding

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:04.762216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:43:03.032611Z digest=sha256:64791a50538bd3797e97305b6fd319bbe4b8463f57f596a67e6f695c2cf4da59

Observation e2174da7-e12b-41df-aeb8-06e665057820 · outbound

This paper cites Chen and William B.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Chen and William B

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:04.741759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:43:03.037293Z digest=sha256:5fa7bab2e492d8d7c9d4d009fef5e8eace5f8f46243d8a12fa6b3017e3ce28ee

Observation 8c0bab9f-0d08-4aef-b355-b32aadb75f55 · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T17:43:03.047959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:43:03.047959Z digest=sha256:b17f44ce260e55578db5db51801e5a59c7cf92901469dda418f9e4177282aeed

Observation 7c0a61f2-f1a3-46bd-9800-de413e0a100f · outbound

This paper cites Gonzalez, Ion Stoica, and Eric P.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Gonzalez, Ion Stoica, and Eric P

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T17:43:03.053816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:43:03.053816Z digest=sha256:842250277bda918df4048015c791bea4771ea285e335cb3de06ab075aac1db31

Observation 53588708-e7e8-44db-86f2-a24f74c2d2c5 · outbound

This paper cites Palm: Scaling language modeling with pathways.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Palm: Scaling language modeling with pathways

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T17:43:03.058756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:43:03.058756Z digest=sha256:771ddce11742d9acb8bd90122aa603881fd1e01183330156dcd9d7d67280af4b

Observation a94e7e06-e704-43d2-8ba4-b05789559123 · outbound

This paper cites an unresolved cited work.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-12T17:43:04.677150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:43:03.063662Z digest=sha256:2f09fd9e32d28b795afc4df8cfec85ece311ed912e9820b9a5e833fddf7fb3bd

Observation 8923c3c4-b5a4-4d60-b4ec-53ed7bae0b2e · outbound

This paper cites EgoQA: Egocentric question an- swering.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding EgoQA: Egocentric question an- swering

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:04.659167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:43:03.068378Z digest=sha256:cdb98ff9fabf0acc2c39ad83d5e8c6eeb6fc8ef364f2d91b316e163fc0357ad3

Observation 5178ef37-024c-4463-8b7d-feb5f9a4adf1 · outbound

This paper cites Study on density peaks clustering based on k-nearest neighbors and principal component analysis.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Study on density peaks clustering based on k-nearest neighbors and principal component analysis

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:04.640474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:43:03.073060Z digest=sha256:bea77f0838a7b526bda804a81792dc7108b27950ee54ed49b6043743ba166523

Observation 114aef3b-3e26-4775-92a2-779f5607bec3 · outbound

This paper cites EV A: exploring the limits of masked visual representation learning at scale.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding EV A: exploring the limits of masked visual representation learning at scale

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:04.622796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:43:03.077586Z digest=sha256:2ed791b07af75f08e79f07cded97564dfd9d2bed63a752dc3378fc11aef8aadd

Observation 7a073534-89e0-4cd5-8319-8f235bfaaaef · outbound

This paper cites Making the V in VQA matter: Elevating the role of image understanding in visual question answering.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Making the V in VQA matter: Elevating the role of image understanding in visual question answering

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:04.604589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:43:03.082185Z digest=sha256:8107bc1fdd5649f7abc8aeba3a35c52c3c1f99949051b6ab3b4a8dce12ed1245

Observation 2b2efb6a-fb1c-44e9-ad25-53a1c8e3cfd0 · outbound

This paper cites Semantic-aware modular capsule routing for visual question answering.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Semantic-aware modular capsule routing for visual question answering

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:04.587504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:43:03.086860Z digest=sha256:6de781c47029c79dba957756c3ee9488580dd1445964a99b19dda7dd5d6c2588

Observation 967e23a5-a606-4c25-8945-23b63da30962 · outbound

This paper cites MA-LMM: memory-augmented large multimodal model for long-term video understanding.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding MA-LMM: memory-augmented large multimodal model for long-term video understanding

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:04.568683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:43:03.091661Z digest=sha256:89ab07e9e2d59098b5a9f755e3daa4f5d5652a02c73173bd77828df6e96ec98d

Observation 5de07298-2e8b-4d31-94ad-f6ed51167c0a · outbound

This paper cites Ma-lmm: Memory-augmented large multimodal model for long-term video understanding.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Ma-lmm: Memory-augmented large multimodal model for long-term video understanding

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:04.551934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:43:03.096166Z digest=sha256:490d515f687a3d63350e9ed1bf23fcdc6582998a70f8846a6d60c9443f0ddde4

Observation 6d4b9996-60c8-438c-b156-f73174c9c75e · outbound

This paper cites Activitynet: A large-scale video bench- mark for human activity understanding.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Activitynet: A large-scale video bench- mark for human activity understanding

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:04.535708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:43:03.100797Z digest=sha256:1c160fbc4cc08a677c901e7be7c973468ba167ac235ba3355de07589d7ac0c4a

Observation 3d217d76-55f0-4919-988e-a7158c81f2f2 · outbound

This paper cites an unresolved cited work.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-12T17:43:04.518297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:43:03.105442Z digest=sha256:4a7be0946dcbebcc86d15d0e733c736325a1f6a77cfea042303f5f35e5e6df5c

Observation 02a8aba9-ed37-42e7-a542-a2bc61b21d20 · outbound

This paper cites VTimeLLM: Empower LLM to Grasp Video Moments.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding VTimeLLM: Empower LLM to Grasp Video Moments

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T17:43:03.110131Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:43:03.110131Z digest=sha256:93bc78f235ed25aeb5ba92c5a87f151e3b715ddd2d87992619148868e8e4133c

Observation a4a3e019-fd4d-42d8-8cec-21765358ad1e · outbound

This paper cites Ng, Hongqiang Rong, and Zichen Li.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Ng, Hongqiang Rong, and Zichen Li

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:04.501170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:43:03.115496Z digest=sha256:3dd383a367c720b2cc5761e03d42eabe5945af1bd9ae786915a6a8a2f137a9a9

Observation 1c862dea-09fb-40c2-8280-c60f3386f8ad · outbound

This paper cites Gqa: A new dataset for real-world visual reasoning and compositional questions.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Gqa: A new dataset for real-world visual reasoning and compositional questions

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:04.484544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:43:03.120120Z digest=sha256:d22754f8bb969f94e429008160f5975f1b82027896ebe00b0a5f81528d2541dc

Observation a63abfb8-567c-49a0-aabe-73c348fdd3a4 · outbound

This paper cites Chat-UniVi: Unified Visual Representation Empowers Large Language Models with Image and Video Understanding.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Chat-UniVi: Unified Visual Representation Empowers Large Language Models with Image and Video Understanding

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T17:43:03.124790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:43:03.124790Z digest=sha256:29bd629076fd4b9c2e4dc711e985caea3792c832f4dc5a9fdebd2bc79fdeee9d

Observation 2611f948-041c-4fcb-8df2-73f54e44b3ed · outbound

This paper cites Scaling Laws for Neural Language Models.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Scaling Laws for Neural Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T17:43:03.130132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:43:03.130132Z digest=sha256:101f0ca19642808dfe29f32947c211c349a0f142d0297c31b34f58ddf147170a

Observation 7b32a324-e24f-4f7b-ac0d-ebe87d4d61ba · outbound

This paper cites an unresolved cited work.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-12T17:43:04.463976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:43:03.135145Z digest=sha256:ebc7d6bd41e05654fd1d0fefba887fef66151be6c234559b90d54ed8f7859960

Observation d1aedadd-e8e9-4b7c-bc4b-b87f8f061e44 · outbound

This paper cites Kingma and Jimmy Ba.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Kingma and Jimmy Ba

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:04.444757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:43:03.140583Z digest=sha256:97f92e83320533e59e1f2efae2750025e810ba2ffa5baacff18231f749100a5f

Observation 1358aabb-24c5-4dfd-84b4-abf6534493aa · outbound

This paper cites Visual genome: Connecting language and vision using crowdsourced dense image annotations.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Visual genome: Connecting language and vision using crowdsourced dense image annotations

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:04.428366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:43:03.145305Z digest=sha256:1bb485e5f78309a195196ebdbc20221cfe69a9cc1682ba999775df94b2f1f3c2

Observation fe1aeb9e-49b2-4b88-8ada-5996889d545e · outbound

This paper cites Visual genome: Connecting language and vision using crowdsourced dense image annotations.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Visual genome: Connecting language and vision using crowdsourced dense image annotations

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:04.411158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:43:03.150322Z digest=sha256:a4d791903b2a8490d30c377ad145886256e51653ac0cdc74bdfe02c62edc7470

Observation 8dd26375-64c0-4b7c-87be-1d82ecc6d9e3 · outbound

This paper cites VideoChat: Chat-Centric Video Understanding.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding VideoChat: Chat-Centric Video Understanding

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T17:43:03.155428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:43:03.155428Z digest=sha256:00bd4c9f5f69fe90fd19a9f5f3eed17b577294123d7f904f204ed58762ca1dd2

Observation 4b36e757-7f2d-47c3-b746-91d1986490f7 · outbound

This paper cites MVBench: A Comprehensive Multi-modal Video Understanding Benchmark.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding MVBench: A Comprehensive Multi-modal Video Understanding Benchmark

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T17:43:03.160467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:43:03.160467Z digest=sha256:a99c957f393a939ce85e01be1641c7458770abd69246743411515c636ca76831

Observation d816dd29-779c-4244-9cf9-18ca9a8829ea · outbound

This paper cites Scienceqa: A new dataset for science question answering.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Scienceqa: A new dataset for science question answering

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:04.392817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:43:03.165529Z digest=sha256:7503a1711ca92acdcea4d95e114ff286ce380eb08e273757f716e4909153fa8e

Observation afc0463e-dede-48a9-9cd3-afd740fe1de1 · outbound

This paper cites Learning dynamic routing for semantic segmentation.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Learning dynamic routing for semantic segmentation

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:04.373424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:43:03.170423Z digest=sha256:ae9cb2e3c3716a81c0c36419a946755fb242c2e60a5a988617aabb13de09ce4f

Observation eb7e9455-6eaf-461f-8f19-563a7e6d1ab0 · outbound

This paper cites LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-12T17:43:03.175123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:43:03.175123Z digest=sha256:7877ced20e530779f8efcc79f4baee4879d4626ecac39812cb562f85f0df32f4

Observation 66d3c174-85f9-447b-a60a-1f540dc777cb · outbound

This paper cites Video-llava: Learning united visual representa- tion by alignment before projection.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Video-llava: Learning united visual representa- tion by alignment before projection

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:04.351160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:43:03.180928Z digest=sha256:88347995c05d4b986c017bbf60b7844ffed41933932b7b3215f9b2760317aab8

Observation c1e97409-d45b-42bf-98a6-7a84d74224b8 · outbound

This paper cites Video-LLaVA: Learning United Visual Representation by Alignment Before Projection.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-12T17:43:03.186135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:43:03.186135Z digest=sha256:8c90836fd8f6718597bcbc46a5ee1ae2c9a002c5fa4612dd5e50e4ae30d9b1c7

Observation 90fcb0c1-438f-463f-8bb4-6fd700a3c9d0 · outbound

This paper cites Vila: On pre-training for visual language models.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Vila: On pre-training for visual language models

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:04.333933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:43:03.191688Z digest=sha256:2933b9956e6aaf9de97af3d1cfc18b1aa438845e7f2d60a7797dd1dd15103e5a

Observation 79bff67f-cb71-4e27-9c1b-2d75451d3a95 · outbound

This paper cites Lawrence Zitnick.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Lawrence Zitnick

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:04.315970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:43:03.196680Z digest=sha256:7ba3e5798be86d6fd395183650745186561271c3e54b270ef159c24c1745072f

Observation 25d71ced-da19-4f2f-a516-afa49035b62f · outbound

This paper cites Improved baselines with visual instruction tuning, 2023.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Improved baselines with visual instruction tuning, 2023

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-12T17:43:03.203557Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:43:03.203557Z digest=sha256:6b5c5ab3305c70097bc074c722b7c23a47fbeb6dd7fc44f870745ee198e307ed

Observation de0d25fb-e774-4205-9732-0b5a640fbf72 · outbound

This paper cites World Model on Million-Length Video And Language With Blockwise RingAttention.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding World Model on Million-Length Video And Language With Blockwise RingAttention

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-12T17:43:03.208853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:43:03.208853Z digest=sha256:005b4fcea51895c14b253312baead715434b4be1bdb6c9ff30598ff00b0aae64

Observation 60a8e0cb-1022-4047-a7ed-825647eab5e4 · outbound

This paper cites BT-Adapter: Video Conversation is Feasible Without Video Instruction Tuning.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding BT-Adapter: Video Conversation is Feasible Without Video Instruction Tuning

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-12T17:43:03.214641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:43:03.214641Z digest=sha256:549dcbe93542c949734cb65b5291c736c56c773859a84524440b537038b0474f

Observation 1593324f-320e-4af4-9cb9-cfc8b1ed7bc7 · outbound

This paper cites ST-LLM: Large Language Models Are Effective Temporal Learners.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding ST-LLM: Large Language Models Are Effective Temporal Learners

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-12T17:43:03.220788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:43:03.220788Z digest=sha256:b5667bf3f242530461d3cf5109dab8ce8bf9cf135a56ae8ea7eef6fa31a77265

Observation 2af014f9-11db-4cdd-aa50-d2a8f3d2c860 · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Learn to explain: Multimodal reasoning via thought chains for science question answering

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-12T17:43:03.226223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:43:03.226223Z digest=sha256:49bb1bba243983c46cffed73718577eb64e15c8190eeac08e0e14fcb34fdcce9

Observation 2ffbb597-19fb-4961-93fd-8e4f31a16a59 · outbound

This paper cites Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-12T17:43:03.231069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:43:03.231069Z digest=sha256:ab91385dad431dd064bc26af992197c7f2635f85cd21600e25615ba9a0730ae0

Observation b8a8055a-73b3-44f0-91dd-8e70d97aa0ee · outbound

This paper cites Some methods for classification and anal- ysis of multivariate observations.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Some methods for classification and anal- ysis of multivariate observations

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:04.272680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:43:03.236284Z digest=sha256:de4d4a38d881eba69f1a0b004a1f09f5db1b960c15a78b293111a296077ec6f9

Observation 1f1405bd-bdfa-4a08-b6e1-179d714a455f · outbound

This paper cites Yuille, and Kevin Murphy.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Yuille, and Kevin Murphy

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:04.253109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:43:03.241058Z digest=sha256:9324926a431098f4a95a47c47178e464ae970643feb7581938e15677a264a3a0

Observation fbd14a40-e224-4093-8831-e1162e9cf647 · outbound

This paper cites Ocr-vqa: Visual question answering by reading text in images.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Ocr-vqa: Visual question answering by reading text in images

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-12T17:43:03.245850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:43:03.245850Z digest=sha256:09b0770cce37da5d2cd82c5ef1c18556e3628655146f18f4f5a5222ca2113814

Observation ba16ab19-d20a-4e06-9572-00797ee7c6ab · outbound

This paper cites Webvidqa: A large-scale dataset for video question answering.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Webvidqa: A large-scale dataset for video question answering

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:04.224590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:43:03.251180Z digest=sha256:c6b8c7ca5346da325e8ccd9b05d14456fe0b980d245a250b3cb0105eabcd64b7

Observation 6ea3099d-e5d3-48e4-be36-b99bf90505b1 · outbound

This paper cites Introducing chatgpt.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Introducing chatgpt

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:04.208955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:43:03.256692Z digest=sha256:6a9e650c91761239d01a6170272a07cd455de321e5370b18e2e8bc38c31331ca

Observation b9e6b080-9048-43f3-8b18-406637cedd42 · outbound

This paper cites GPT-4o system card, 2024.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding GPT-4o system card, 2024

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:04.192873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:43:03.261844Z digest=sha256:375e83737704801bfa960ccfbf36d15e8a2787848d9141923c8392845347e49d

Observation 6ecae718-bdbf-4511-a75c-0ff1b842447e · outbound

This paper cites Learning transferable visual models from natural language supervision.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Learning transferable visual models from natural language supervision

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:04.177914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:43:03.267235Z digest=sha256:4b9200154b3d09b77f2e47c6c15b5c74c0e49e9ab32e433de1f79029de44d685

Observation b2a3c4cd-944a-46b1-a257-bab491b451bf · outbound

This paper cites Improving language understanding by gener- ative pre-training.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Improving language understanding by gener- ative pre-training

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:04.162143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:43:03.272111Z digest=sha256:ee7867105cdce3debc484a60b1367cbe178b47bdac7ce40794b8906fcaf10752

Observation 80ab209e-d25d-468f-b4f5-f2c99e816a88 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:04.146210Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:43:03.277089Z digest=sha256:2138aec9178559805992d8234af292010e9a83c958bed5bde7aacb73a0c56344

Observation 849ce177-89ab-4c57-84b2-263947569873 · outbound

This paper cites Timechat: A time-sensitive multimodal large language model for long video understanding.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Timechat: A time-sensitive multimodal large language model for long video understanding

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:04.129764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:43:03.281962Z digest=sha256:2e8bd0b6c102e0887f4589bbec9290a6c986dac5eca0127114a24a8185e9855b

Observation c8db052f-139d-4721-92db-375d350333a0 · outbound

This paper cites Massof Sarah L.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Massof Sarah L

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:04.112583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:43:03.288087Z digest=sha256:cb34b4c1dc71058703a78a4afc81fdfcfd08307cdab3040db30693f2877d957b

Observation e74eb7fb-e81a-4a1a-aedb-f5c5f9d1a06d · outbound

This paper cites A-okvqa: A bench- mark for visual question answering using world knowledge.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding A-okvqa: A bench- mark for visual question answering using world knowledge

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:04.095257Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:43:03.292697Z digest=sha256:8540432166d165a8f0c7d880dd2a7fb87ef0b72b137a4e7c6416f7ab760940c4

Observation 35c00115-a116-4de1-9640-eabf185e24a1 · outbound

This paper cites Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:04.078763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:43:03.297505Z digest=sha256:a162aac9d7164ec48510f42130a1511245abe6845ec6ec44288dd7f574cc2df1

Observation 04cbea46-a537-4ccb-a7c1-a20ce0943b2b · outbound

This paper cites Textcaps: a dataset for image captioning with reading comprehension.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Textcaps: a dataset for image captioning with reading comprehension

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:04.059597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:43:03.302047Z digest=sha256:c39a53702d2269e5f7e4ff10f82b259fa98fff2261db5a9ccc3b8367d373a284

Observation 07fcd0e9-4a9d-48f4-913e-16a82ba77ff9 · outbound

This paper cites MovieChat: From Dense Token to Sparse Memory for Long Video Understanding.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding MovieChat: From Dense Token to Sparse Memory for Long Video Understanding

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-12T17:43:03.306611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:43:03.306611Z digest=sha256:c14e3e6db54cab5f2b818709722466c4ed0b79aa08182f06c03fce58c3df5063

Observation 244acf83-61ea-4a6b-9dd9-53248aa8ebca · outbound

This paper cites Moviechat: From dense token to sparse memory for long video understanding.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Moviechat: From dense token to sparse memory for long video understanding

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:04.038794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:43:03.311747Z digest=sha256:7bd46da2dba4afb2c5432060162ca62f848ccf4976bb12540fa68a067cd539bc

Observation adc1f000-41c0-4e15-beea-8be646f352d0 · outbound

This paper cites DreamFrame: Enhancing Video Understanding via Automatically Generated QA and Style-Consistent Keyframes.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding DreamFrame: Enhancing Video Understanding via Automatically Generated QA and Style-Consistent Keyframes

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-12T17:43:03.316445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:43:03.316445Z digest=sha256:fe5c9fe76fb6120600da86abce516585be4690ff853f88453bb3cd03ad33b92b

Observation f27cfd0f-7d32-4381-9795-91a3214f24e6 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding LLaMA: Open and Efficient Foundation Language Models

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-12T17:43:03.321491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:43:03.321491Z digest=sha256:f63fe3260e7303519490cb385a55b520e0ba67c7bab7bf7ac0531c21f630c790

Observation 43a75b67-c0a8-435b-9aac-d5c43325d460 · outbound

This paper cites Ocrvqa: A new dataset for optical character recognition in visual question answering.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Ocrvqa: A new dataset for optical character recognition in visual question answering

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:04.020499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:43:03.326429Z digest=sha256:61686b59ef9142dbbd84e74b0ad1a69954e76438c950aff0d7dacb35c8af0306

Observation 7e892c22-933f-4318-965b-dc19ed6cf8a9 · outbound

This paper cites LVBench: An Extreme Long Video Understanding Benchmark.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding LVBench: An Extreme Long Video Understanding Benchmark

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-12T17:43:03.331072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:43:03.331072Z digest=sha256:0d6f29f75cb6366f7b869edbad2fb108e96dc25ae01d55f600bcb657206a7350

Observation 97f740c3-73c6-4c9e-8958-67f8bdfdf36e · outbound

This paper cites VideoHallucer: Evaluating Intrinsic and Extrinsic Hallucinations in Large Video-Language Models.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding VideoHallucer: Evaluating Intrinsic and Extrinsic Hallucinations in Large Video-Language Models

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-12T17:43:03.336459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:43:03.336459Z digest=sha256:7a8dedb7bb96f9bcea3e38e8a9d545f1f75511c01188d9a8ab02b68062edba93

Observation bd29c8d1-fd60-4a56-b0ad-e2dc3962211b · outbound

This paper cites Videollamb: Long video understanding with recurrent mem- ory bridges.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Videollamb: Long video understanding with recurrent mem- ory bridges

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:04.003342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:43:03.341593Z digest=sha256:6a79fbc93212d7e74eb8fe957a3b8320c60bb533064ef5579ee4b6fc86770cea

Observation cc9702e0-bbdb-40bd-a4b7-33a4e5b83220 · outbound

This paper cites FreeVA: Offline MLLM as Training-Free Video Assistant.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding FreeVA: Offline MLLM as Training-Free Video Assistant

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-12T17:43:03.346619Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:43:03.346619Z digest=sha256:31518b86c0012798a51a79a25d3b69538539f9b90ea2a1ae9c86eda4773755ba

Observation 0d789eab-03a8-4918-a2c7-8a4e8d87ea3a · outbound

This paper cites Davis, Kristen Grauman, and Rog´erio Schmidt Feris.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Davis, Kristen Grauman, and Rog´erio Schmidt Feris

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:03.986333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:43:03.352245Z digest=sha256:06bc5b30b9889f8d9c8806a0daa5fe8e183c1b47ade13f26aa6981d4fba2ed54

Observation c824db95-0761-4a49-8aef-c606a9170226 · outbound

This paper cites Deep learning for video classification and captioning.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Deep learning for video classification and captioning

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:03.969334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:43:03.356916Z digest=sha256:cb2ea7eea1b7b7dc4efdc3ef41ecfbf521526d940d4021c9edb95163ddd0181e

Observation b0b61000-7204-46e3-9c89-70cbe5eb90f7 · outbound

This paper cites Msr-vtt: A large video description dataset for bridging video and language.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Msr-vtt: A large video description dataset for bridging video and language

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-12T17:43:03.361819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:43:03.361819Z digest=sha256:e52945ab66a3b42a52930f723eed6ec3d49f40058f1df3eb7a4723c1324a1aca

Observation b417a12e-c613-4623-b461-5dcfb2fc48c0 · outbound

This paper cites MSR-VTT: A large video description dataset for bridging video and language.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding MSR-VTT: A large video description dataset for bridging video and language

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:03.940819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:43:03.366481Z digest=sha256:3995b65d0a6a77740364c15408dbbd412f7682f2c7bd0588363c3e86399eb9c1

Observation 580d49b8-86fd-4237-b6e6-0aaf0f50a8b1 · outbound

This paper cites PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-12T17:43:03.371221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:43:03.371221Z digest=sha256:33f9faeddc04ec9968d97045a3d2593cbacd89eb890d07420e98395009bcce1b

Observation 0886b48f-09a5-4d1f-a63a-4601f2da034d · outbound

This paper cites CLEVRER: CoLlision Events for Video REpresentation and Reasoning.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding CLEVRER: CoLlision Events for Video REpresentation and Reasoning

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-12T17:43:03.376477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:43:03.376477Z digest=sha256:606a8ebfd4faf877182dc93ae83b7c3bf082d63d389b01a133f23ae1a644474e

Observation 86ed1b74-7c25-43e6-9888-5a20b69ec836 · outbound

This paper cites Tenenbaum.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Tenenbaum

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:03.924319Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:43:03.381552Z digest=sha256:2a7186e2e3be203928234d031a597b62326243b5d66aeebb20ac6a7a06632699

Observation ac5110db-7965-434d-a3c8-be105974b135 · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-12T17:43:03.386720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:43:03.386720Z digest=sha256:9cf6e25095a6cdce9f08c9c329e721368f8ef0593a1e0f2372e7feae562cc617

Observation 5cbfdf1e-ed6a-4579-8036-f4c7b9d9ea87 · outbound

This paper cites Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-12T17:43:03.391576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:43:03.391576Z digest=sha256:31456ed7eac4261b59a6f544801dff8d280ff545a2d587ffbe26e4795a88bceb

Observation f67af649-7e95-4e51-9057-9c8392cd8a03 · outbound

This paper cites Llama- adapter: Efficient fine-tuning of language models with zero- init attention.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Llama- adapter: Efficient fine-tuning of language models with zero- init attention

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:43:03.906623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:43:03.396276Z digest=sha256:6f10a7e0a1c628c28c7df906a5c8e533b4828af6965edfa13f662703001583b6

Observation f7d1270e-acba-4a7c-b2a9-061431d94cc9 · outbound

This paper cites Please Carefully Think.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Please Carefully Think

Reference 79

Resolution
malformed identifier
raw_fallback, observed 2026-08-12T17:43:03.887845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:43:03.400998Z digest=sha256:dc78ca373219731b255b95a902a17ade1ea07eaac16e1aa4c80f53f95215a8eb

Observation 831ab334-8259-405f-a2cb-5c78bfdc1331 · outbound

This paper cites an unresolved cited work.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding Unresolved cited work

Reference 200

Resolution
unresolved
raw_fallback, observed 2026-08-12T17:43:04.719668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T17:43:03.042499Z digest=sha256:a11e96ffcf05d4f0c19aaf9083c04fd82ef4a0cb69af3cc74a68f3294937c642

Pith citing papers

No inbound Pith citation observations are available.