Pith. sign in

Paper Citation Record · LEDGER

DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding

As of 10 August 2026, this Paper Citation Record lists 57 of 57 outbound references and 0 inbound Pith citation observations for arXiv:2507.15569.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.15569 v1

Coverage vector

measured 57 of 57 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T15:34:44.512500Z

measured 57 of 57 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

57 of 57 outbound references displayed

  • verified exact0
  • verified fuzzy32
  • unresolved25
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f2aa759d-540a-4dd3-9b35-fcdbadc800c3 · outbound

This paper cites Flamingo: a visual language model for few-shot learning.NeurIPS, 2022.

DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding Flamingo: a visual language model for few-shot learning.NeurIPS, 2022

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T15:34:44.354775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:34:44.354775Z digest=sha256:d2328cf4d813659a76c24f4b3089b65aaf5f4c4a552eb036247853efe913bdc6

Observation f12f1706-f49e-452d-bde1-3937039bfe13 · outbound

This paper cites Vivit: A video vision transformer.

DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding Vivit: A video vision transformer

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T15:34:44.358353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:34:44.358353Z digest=sha256:d67b60dbd51a080559722b543c7e5d5402520573685c98b4c50f710e9fa66e52

Observation a810cc3f-b520-4e8b-837c-051629199850 · outbound

This paper cites Exploring Visual Prompts for Adapting Large-Scale Models.

DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding Exploring Visual Prompts for Adapting Large-Scale Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T15:34:44.361412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:34:44.361412Z digest=sha256:e6249b18155534bc10136c4a6b248c1b0dd69134e2a2528ddc7ed81e552a5d8c

Observation 9204a1de-db76-4350-988f-6541a262e604 · outbound

This paper cites Relevant intrinsic feature enhancement network for few-shot semantic segmentation.

DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding Relevant intrinsic feature enhancement network for few-shot semantic segmentation

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:34:45.186177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T15:34:44.364929Z digest=sha256:50c7c46bae3b07de58386a53968600d21e70ff64b1a45f8b3f064d540c9494d6

Observation d902335b-31c9-4ce7-b95e-cd24b9314925 · outbound

This paper cites Cores: Orchestrating the dance of reasoning and seg- mentation.

DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding Cores: Orchestrating the dance of reasoning and seg- mentation

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:34:45.177037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T15:34:44.368268Z digest=sha256:50ff09c75662e218b90c2ab2ea24e1f8f3bbf75d769806735222502db1974ec4

Observation 97a9d4ab-5612-42dc-9568-a85eafd9d064 · outbound

This paper cites Is space-time attention all you need for video understanding? In ICML, page 4, 2021.

DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding Is space-time attention all you need for video understanding? In ICML, page 4, 2021

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T15:34:44.371708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:34:44.371708Z digest=sha256:fa992908f4c3c2a67a5739ee9c2e77ffa637cb8ee47058b1683f2da78d9f7662

Observation ed27e943-8da2-48a9-b119-bd9c8fb48b69 · outbound

This paper cites Activitynet: A large-scale video benchmark for human activity understanding.

DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding Activitynet: A large-scale video benchmark for human activity understanding

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:34:45.161693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T15:34:44.374741Z digest=sha256:4fb499af4357c1e9fc5d3d84a3a5b754a0671c1bfd8f826bff4586e680ca4864

Observation 047a7701-6c38-4280-b268-a02799f653c6 · outbound

This paper cites Quo vadis, action recognition? a new model and the kinetics dataset.

DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding Quo vadis, action recognition? a new model and the kinetics dataset

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:34:45.151912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T15:34:44.377884Z digest=sha256:7734d43cac4f6e48d0386e1e9364a645c630e5f597c4c4280d56408f26659182

Observation 31550c44-a080-4788-aebf-c472fb184c73 · outbound

This paper cites Deep temporal linear encoding networks.

DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding Deep temporal linear encoding networks

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:34:45.142530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T15:34:44.380704Z digest=sha256:b51ca398959377aec805bad50e992f4debb6053474da814b69ffa8e67cb88e0d

Observation 01a65432-7fe1-4d89-abb6-9290af5a52c7 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T15:34:44.383416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:34:44.383416Z digest=sha256:13f77be9eeee0d1cef8c9d0a352903ab2cd6daf187ffc7341685b4b64e5ef241

Observation efca8d38-954e-475a-8066-dd34ff2a7bde · outbound

This paper cites Convolutional two-stream network fusion for video action recognition.

DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding Convolutional two-stream network fusion for video action recognition

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:34:45.131871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T15:34:44.386251Z digest=sha256:0513749733ece16a0492cc832deb09fa8d956ca7bd2400208e8b11e0d6b0b3ce

Observation 06b216d4-bf76-415d-bfec-5b4821cd3b8e · outbound

This paper cites Slowfast networks for video recognition.

DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding Slowfast networks for video recognition

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T15:34:44.388931Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:34:44.388931Z digest=sha256:7bca6b6cad424199433ef13b377d9ca31b32f5bc538c7e64de760f7d734e6805

Observation fe50a836-005b-4f99-9b63-5f8df3e74e80 · outbound

This paper cites A Systematic Survey of Prompt Engineering on Vision-Language Foundation Models.

DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding A Systematic Survey of Prompt Engineering on Vision-Language Foundation Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T15:34:44.391857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:34:44.391857Z digest=sha256:27abc3cf4bc348729bb3c54e2b5a708c767a52175b3f4f6f2d15bba5e5c68716

Observation 6c83cb25-e520-4df3-9bd8-c0d822c1f4a9 · outbound

This paper cites Lita: Language instructed temporal-localization assistant.

DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding Lita: Language instructed temporal-localization assistant

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:34:45.115808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T15:34:44.394727Z digest=sha256:e4265a70dec651c686a931c937eb1701380c56df3ef0d6fa8acd2479255cf87b

Observation 8758f006-c3e4-4eaf-bf76-4b97f6f07d39 · outbound

This paper cites Vi- sual prompt tuning.

DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding Vi- sual prompt tuning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T15:34:44.397334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:34:44.397334Z digest=sha256:8146a7420873b9df4bb05139aa01607820d389382df9cf7dded592681c019625

Observation 76e16600-94a2-4da5-a431-b8651859cbbc · outbound

This paper cites Chat-univi: Unified visual representation em- powers large language models with image and video un- derstanding.

DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding Chat-univi: Unified visual representation em- powers large language models with image and video un- derstanding

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:34:45.099303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T15:34:44.400058Z digest=sha256:9850278fc4767b1c98795836d942e7a01a17d2df4255bae7fcc7b36a4243b5f4

Observation 6b2d06ef-f871-479b-bf41-758ceadb3ae8 · outbound

This paper cites Video-lavit: Unified video-language pre-training with decoupled visual-motional tokenization.

DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding Video-lavit: Unified video-language pre-training with decoupled visual-motional tokenization

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T15:34:44.402876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:34:44.402876Z digest=sha256:44f67b0b9e7b5e4e7d3d931734be5a36e382db323d730f4e8ff44694993ae667

Observation 6efde5e5-e056-4173-8d90-cf61eefb4bd5 · outbound

This paper cites Large-scale video classification with convolutional neural networks.

DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding Large-scale video classification with convolutional neural networks

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:34:45.089195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T15:34:44.405446Z digest=sha256:d3b90033e7c0b212e56f355a176f2baada829278ba1f70ff8b6b2b03995451e5

Observation 886d5255-bf2c-432e-9ee2-bafee1ce1371 · outbound

This paper cites Maple: Multi-modal prompt learning.

DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding Maple: Multi-modal prompt learning

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:34:45.077119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T15:34:44.407980Z digest=sha256:937913006226a2ffbee281bea54ce5fc5ee9c49d3ab216ae8d6647ca1b9c376b

Observation 86d87ae0-451a-468d-ba72-f19c6f13db1c · outbound

This paper cites An Image Grid Can Be Worth a Video: Zero-shot Video Question Answering Using a VLM.

DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding An Image Grid Can Be Worth a Video: Zero-shot Video Question Answering Using a VLM

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T15:34:44.410539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:34:44.410539Z digest=sha256:997c0b8d62192009f2d93386927b92586b75714875054a4c04e86ba9ce64f477

Observation 78686686-890a-4340-87c9-c72b035b5a40 · outbound

This paper cites LISA: Reasoning Segmentation via Large Language Model.

DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding LISA: Reasoning Segmentation via Large Language Model

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T15:34:44.413798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:34:44.413798Z digest=sha256:813f42f816fa9b0cbade74603b1248b650a59bca0ff2cec7a20aa8b5cd094f73

Observation e67cff02-1234-4df9-85db-504bce9b20d6 · outbound

This paper cites Deep local video feature for action recog- nition.

DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding Deep local video feature for action recog- nition

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:34:45.067593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T15:34:44.416913Z digest=sha256:2988fcba27f52fa528a90f5266d196b97d4a49d005e47adfff516224d53f0679

Observation cbe3b580-50c7-4eb6-9bc5-41608f5d5610 · outbound

This paper cites BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models.

DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T15:34:44.419656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:34:44.419656Z digest=sha256:daefa916396c9fe1d0378d6f0bde73cc5e365231293689100de92a749724bb87

Observation 1fe9950b-56f2-4ab2-b4a8-f37eaedb4cb7 · outbound

This paper cites VideoChat: Chat-Centric Video Understanding.

DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding VideoChat: Chat-Centric Video Understanding

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T15:34:44.422636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:34:44.422636Z digest=sha256:1a57a53743e46a2b2128a2bcbeb6fedfbefd32c554a4c8683cb7879c14a9282f

Observation 3906ca21-74a7-436b-9fba-b566579c0520 · outbound

This paper cites Mvbench: A comprehensive multi-modal video understand- ing benchmark.

DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding Mvbench: A comprehensive multi-modal video understand- ing benchmark

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:34:45.057513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T15:34:44.425250Z digest=sha256:341e07eb1bc0c2dc04d42979a60dc223e532b88ab49cc16a499e055f2ea85940

Observation 4f78f31d-0889-4c27-ab81-189ccb23a05e · outbound

This paper cites Tgif: A new dataset and benchmark on animated gif description.

DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding Tgif: A new dataset and benchmark on animated gif description

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:34:45.048001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T15:34:44.427983Z digest=sha256:134e8b435dd9c8fc158ad7ca45616b3058606afd43bd0569be933a65eed4a97a

Observation 5d3997aa-1479-4679-a207-0baac91a52ba · outbound

This paper cites Llama-vid: An image is worth 2 tokens in large language models.

DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding Llama-vid: An image is worth 2 tokens in large language models

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:34:45.038073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T15:34:44.430570Z digest=sha256:888dbfd456461cf3725bc00d3ae7f49cefad58c0940c022b068c22f230410781

Observation 927376e1-c046-4e7c-9352-a6564ff8a61d · outbound

This paper cites Video-LLaVA: Learning United Visual Representation by Alignment Before Projection.

DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T15:34:44.433210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:34:44.433210Z digest=sha256:0e7bb566fe68ec11734f0745ecaf836370aa796ccd8497144c04715fe08ade06

Observation 42037cba-cf27-4151-b814-ad705284f57f · outbound

This paper cites Visual instruction tuning.

DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding Visual instruction tuning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T15:34:44.436170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:34:44.436170Z digest=sha256:4065441de649add7205fa0080631e4c01306e721f86755488f88b0c9e9367dad

Observation f149dbf1-c940-4306-ba9b-dcb387911ee1 · outbound

This paper cites Pre-train, prompt, and predict: A systematic survey of prompting methods in nat- ural language processing.

DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding Pre-train, prompt, and predict: A systematic survey of prompting methods in nat- ural language processing

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:34:45.022210Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T15:34:44.438941Z digest=sha256:ef9d2594980571394636acffdf609907f940fdcd51895d1ccabab868e6843a6b

Observation f949c9ea-8c37-4d96-aab7-71637291e73b · outbound

This paper cites St-llm: Large language models are effective tem- poral learners.

DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding St-llm: Large language models are effective tem- poral learners

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:34:45.011156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T15:34:44.441777Z digest=sha256:7f4b5b0bebc036d2b63b3c723e4f4363997ab51cbe809c1274ed5f56473f9929

Observation 0a00f510-85b7-4864-9be1-0d8183088207 · outbound

This paper cites PPLLaVA: Varied Video Sequence Understanding With Prompt Guidance.

DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding PPLLaVA: Varied Video Sequence Understanding With Prompt Guidance

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T15:34:44.444553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:34:44.444553Z digest=sha256:0ef973ccb3c6daade4f04c65fcff48d7226846da2c513d91ca051bbf1365239e

Observation 25f1e70a-1b36-485c-9ad4-bcbe0a679412 · outbound

This paper cites Hybrid-level instruction injection for video token com- pression in multi-modal large language models.

DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding Hybrid-level instruction injection for video token com- pression in multi-modal large language models

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:34:45.001920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T15:34:44.447387Z digest=sha256:966ce004cd0d1d998f7eb73c944ae11c7a7f9c28d343f47fba624d9a680d880f

Observation 22f7bdd3-7318-42a1-9806-3829e0bcfd82 · outbound

This paper cites Vista-LLaMA: Reducing Hallucination in Video Language Models via Equal Distance to Visual Tokens.

DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding Vista-LLaMA: Reducing Hallucination in Video Language Models via Equal Distance to Visual Tokens

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T15:34:44.450026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:34:44.450026Z digest=sha256:3481a6cc299209d0a41410f5384f474595a29d9a8c4fe7e87cba8b67f1c9f5cb

Observation b074dd2b-5d0c-4053-9b0b-bcacb9f59593 · outbound

This paper cites Video-chatgpt: Towards detailed video understanding via large vision and language models.

DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding Video-chatgpt: Towards detailed video understanding via large vision and language models

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:34:44.992531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T15:34:44.453154Z digest=sha256:d6ae68405c65cd5c0d86729d9b5606806725366a8c2d2557dcf91e7a0d280889

Observation a26b246b-3b7a-4459-9e97-76a92169cba1 · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding Learning transferable visual models from natural language supervi- sion

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:34:44.982375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T15:34:44.456204Z digest=sha256:9e4d8e440b4384c816065da4d230de903b6b42221224b48bbf6760a7475febcf

Observation 0b4c4c9b-ef73-4f35-a271-363c8546f5fe · outbound

This paper cites Mul- titask vision-language prompt tuning.

DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding Mul- titask vision-language prompt tuning

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:34:44.971774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T15:34:44.458932Z digest=sha256:1bbccca793600c0ab014e6b02dbf959b1b0c956a2294f96aefa586cd0610954e

Observation a007cafb-bd39-468d-850d-151dc2858a56 · outbound

This paper cites What does clip know about a red circle? vi- sual prompt engineering for vlms.

DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding What does clip know about a red circle? vi- sual prompt engineering for vlms

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:34:44.962147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T15:34:44.461665Z digest=sha256:0cf32f19f04e3097928bcfbc328cbe475cbc9dde111014a14166f86ad54cad9d

Observation 0b9390b3-047e-40e2-885b-53296c9df236 · outbound

This paper cites Two-stream con- volutional networks for action recognition in videos.

DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding Two-stream con- volutional networks for action recognition in videos

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:34:44.952584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T15:34:44.464208Z digest=sha256:2a71a696c1b5a9b5e4019d10dcb0197fd498d4ea7bdb499cd3044eafa1d1e290

Observation 8392e049-9274-45a9-a852-e6fb520459f0 · outbound

This paper cites Moviechat: From dense token to sparse memory for long video understanding.

DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding Moviechat: From dense token to sparse memory for long video understanding

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:34:44.943100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T15:34:44.466804Z digest=sha256:c6721351f69dc453339397506c19b14660c6bdf3bef0d0b6fa197642110ec5f9

Observation 5339d259-49d9-44ea-833b-97d2c51089c3 · outbound

This paper cites Ufo: A unified approach to fine-grained visual perception via open- ended language interface.

DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding Ufo: A unified approach to fine-grained visual perception via open- ended language interface

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T15:34:44.469414Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:34:44.469414Z digest=sha256:0bfdc4e3331324d16f54182c60be234eaf6f9bae419602398d7cd9545131dfe3

Observation f116e752-67e6-49a2-b634-1073b17f5477 · outbound

This paper cites Qwen2.5: A party of foundation models, 2024.

DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding Qwen2.5: A party of foundation models, 2024

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:34:44.932971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T15:34:44.472057Z digest=sha256:a04e444816d41c04d884dd8d312017be4a5642172b8867efd44041b639d27382

Observation 90ac6e1f-5dd4-4f91-a9d4-2e4a9e11e204 · outbound

This paper cites Learning spatiotemporal features with 3d convolutional networks.

DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding Learning spatiotemporal features with 3d convolutional networks

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T15:34:44.474572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:34:44.474572Z digest=sha256:68bf6caccfe9291ebe4a5a3844ae70b1eb2426383c244d667176f8b4fdac1b10

Observation bf7cea43-2d1b-484f-8642-ab3d228a0e81 · outbound

This paper cites A closer look at spatiotemporal convolutions for action recognition.

DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding A closer look at spatiotemporal convolutions for action recognition

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:34:44.917281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T15:34:44.477334Z digest=sha256:4f878d92a0ab6ee1ede5267f23600fb502f52a4accf99b78ab7ff76d6cf10294

Observation 562445aa-e6ce-488a-ac70-6da7fe1054ed · outbound

This paper cites Action recogni- tion with trajectory-pooled deep-convolutional descriptors.

DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding Action recogni- tion with trajectory-pooled deep-convolutional descriptors

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:34:44.907511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T15:34:44.480128Z digest=sha256:a75b554b9dba905f823c35c97cbb35dbd3b4f13cba53bf3c72524bb821b8f2c5

Observation 1ac8883f-6190-43e9-9dd7-56414c619435 · outbound

This paper cites Temporal segment net- works: Towards good practices for deep action recognition.

DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding Temporal segment net- works: Towards good practices for deep action recognition

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:34:44.898103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T15:34:44.482802Z digest=sha256:76b3c5fe030ae5331711cf3ed4af13f90223b85ab3cedcb4aecdf9207fd2d9c6

Observation 2d08469b-7499-4da5-9a0c-291c211afaf6 · outbound

This paper cites Deep learning for video classification and captioning.

DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding Deep learning for video classification and captioning

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:34:44.887994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T15:34:44.485388Z digest=sha256:dd68f930acd35739fa71113eed54fc333b3a10b0a8adbe1597cc9fe8178c5d4b

Observation 04648ed4-742e-4fc7-a6dc-0c4f81a1dd1e · outbound

This paper cites Msr-vtt: A large video description dataset for bridging video and language.

DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding Msr-vtt: A large video description dataset for bridging video and language

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:34:44.878270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T15:34:44.488052Z digest=sha256:7069e63ef31813842bee83a17882600e1ce67bc0bccc2e495d6a185e24b57675

Observation 62968b3d-c61b-44b4-b4b3-72671e84ee31 · outbound

This paper cites PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning.

DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T15:34:44.490784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:34:44.490784Z digest=sha256:e311d2389d9b63e5199dc196cdf32e4f0b359106111c02f3ae74820cdb5daf53

Observation 39de8a1a-26a0-4798-a23c-7dd48455c721 · outbound

This paper cites Beyond short snippets: Deep networks for video classification.

DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding Beyond short snippets: Deep networks for video classification

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:34:44.868532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T15:34:44.493641Z digest=sha256:5ed5dbde58ff6dda3b0c3c5f6f0046039a3515cfad0e013853efca04165a1e4b

Observation 624eca69-ece9-48ad-a3fe-6ee505763a9a · outbound

This paper cites Unified Vision and Language Prompt Learning.

DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding Unified Vision and Language Prompt Learning

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T15:34:44.496621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:34:44.496621Z digest=sha256:66c0bba81bf211a109ff05409696723f434f0bec71f1f890841b288d58bf2afe

Observation 4e413645-f57d-425e-b8af-00f7ad24a065 · outbound

This paper cites Sigmoid loss for language image pre-training.

DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding Sigmoid loss for language image pre-training

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:34:44.858675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T15:34:44.499401Z digest=sha256:c6edf97049297ebcb3d2bcae62e1cd80dae4163628df9ef6be55e86fbed27cb1

Observation cadc5b69-4b8b-4b22-94c9-89979ec767a7 · outbound

This paper cites Real-time action recognition with enhanced motion vector cnns.

DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding Real-time action recognition with enhanced motion vector cnns

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:34:44.848526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T15:34:44.502074Z digest=sha256:e86e82cb48d8515c2f679d2303a88170c0d13a01089889569ca29c01a03dc4e7

Observation 5bb3d3fb-30bd-42fc-8a91-94d8c352daf4 · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T15:34:44.504615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:34:44.504615Z digest=sha256:c1191901f4f13d6b87a25846b5066ca9decd7c27026143995d0109477129200e

Observation 2e549f92-9770-44e1-8e79-718f53c45890 · outbound

This paper cites Conditional prompt learning for vision-language mod- els.

DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding Conditional prompt learning for vision-language mod- els

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T15:34:44.507428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:34:44.507428Z digest=sha256:b01301b4e2f090b1045f4cc122f21e0f3eb25a38c183a43bbabd6111f4b5f520

Observation 3e308a2b-6870-4be2-8fd7-30105dfceff0 · outbound

This paper cites Learning to prompt for vision-language models.

DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding Learning to prompt for vision-language models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T15:34:44.510044Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:34:44.510044Z digest=sha256:715ace8823c5bbe4574f079af87e54f4b51c4504cb4fb93e78f1101bfde4f83b

Observation 35094d80-6c56-4e63-8554-ee5219e9d0da · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-06T15:34:44.512500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:34:44.512500Z digest=sha256:d45fc2fb7052589792af8e4d308704bdc4c3489917c6d4248cc5ca3e3d7c42c7

Pith citing papers

No inbound Pith citation observations are available.