Pith. sign in

Paper Citation Record · LEDGER

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders

As of 20 August 2026, this Paper Citation Record lists 84 of 84 outbound references and 3 inbound Pith citation observations for arXiv:2505.24158.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.24158 v1

Coverage vector

measured 84 of 84 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:37:24.588694Z

measured 87 of 87 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T14:39:08.699220Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-12T03:06:18.765685Z

Reference resolution

84 of 84 outbound references displayed

  • verified exact1
  • verified fuzzy52
  • unresolved30
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0b0f98ab-5ecc-4387-99bc-7583abcaa79a · outbound

This paper cites GPT-4 Technical Report.

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:16.067831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:16.067831Z digest=sha256:661f48a4be997c8077fac2ca1093b67c713c45078034bc135b8e96c04bf380a6

Observation 2a25c717-9def-4308-8ff1-c6049a818e78 · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Flamingo: a visual language model for few-shot learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:16.151840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:16.151840Z digest=sha256:662b390c60a5a78f0f2778db30e96ea9b9f54520e03a9e6ef50be4380751d1f8

Observation 68155e93-d9bf-49c6-b5da-6cad2496cd05 · outbound

This paper cites Vqa: Visual question answering.

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Vqa: Visual question answering

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:16.236696Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:16.236696Z digest=sha256:d347e738c4b33eb2ebb8611ca9b41e80c4fd11b3b340d3c99a090ef7cfa9fc41

Observation 5bc3e1f1-f3a4-4f92-a82e-04cd9341d2bc · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:16.323580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:16.323580Z digest=sha256:fe1cb8d5c9c7d9a6612c3c0d6a12464ca7da7d3206af5d69870d376c0238c51c

Observation e7c4e58b-8675-472d-a7ed-ac3f59df9c45 · outbound

This paper cites Solving mixed-integer quadratic programming problems with ibm-cplex: a progress report.

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Solving mixed-integer quadratic programming problems with ibm-cplex: a progress report

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:16.421314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:16.421314Z digest=sha256:d42ad5721cc6eb094284d28060814bcdc41c4dfa5614e5f0c0771067198b2012

Observation 02f8b0ea-6d17-4e2d-9aba-f79e9be26d52 · outbound

This paper cites On the Opportunities and Risks of Foundation Models.

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders On the Opportunities and Risks of Foundation Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:16.506733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:16.506733Z digest=sha256:ccc556a480ae74d8b27895fd30389ab665b1253f192ab7da0305066463c78b6d

Observation 8b1037ce-82e4-47c4-a89d-2e18f8978df0 · outbound

This paper cites Language models are few-shot learners.

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Language models are few-shot learners

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:16.620624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:16.620624Z digest=sha256:5e531b129b717005fbbd4d7893b222a6d994be7b586f4fa22f147ca13a03994d

Observation 896ac7d1-7c57-4717-a50c-28a592b161b9 · outbound

This paper cites Hourvideo: 1-hour video-language understanding.

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Hourvideo: 1-hour video-language understanding

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:37:34.498288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:37:16.681958Z digest=sha256:b0c9cfcc18ff894de0c7527502ab49ece25664e202a46203852d128892fdd2c1

Observation c3568a3a-b413-41e3-88e1-e0f954ffb8ae · outbound

This paper cites Sharegpt4video: Improving video understanding and generation with better captions.

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Sharegpt4video: Improving video understanding and generation with better captions

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:37:34.329297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:37:16.776741Z digest=sha256:e16f7928c5cc4995a2d15f5cbb7b56cc9f6506e97c334a2e588bf4dd3745d8ed

Observation 27e4ed29-6118-42b5-b76a-559978453cd2 · outbound

This paper cites LongVILA: Scaling Long-Context Visual Language Models for Long Videos.

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders LongVILA: Scaling Long-Context Visual Language Models for Long Videos

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:16.941581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:16.941581Z digest=sha256:3381b869fd4eb55df461b0a8dc900f3ae5d5488191bcd93e1ef7d775c1393b0f

Observation e3b44931-b510-4d07-8757-d9f2446ea122 · outbound

This paper cites Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks.

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:37:34.139239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:37:17.185352Z digest=sha256:dbe6488b83344fd466c511a23c8478f72887abd2d27f24171f96c9f63fd28022

Observation 762e141e-19cd-44e0-bc51-3a40d50f8291 · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:17.344219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:17.344219Z digest=sha256:79f5222f2433a6e572b8f753d117a0cc42b1d3fa59bb8e3c2dcbe0e66a8177e2

Observation d484031f-54be-415b-940f-bb6fb78340da · outbound

This paper cites Patch n’pack: Navit, a vision transformer for any aspect ratio and resolution.

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Patch n’pack: Navit, a vision transformer for any aspect ratio and resolution

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:37:33.935943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:37:17.455893Z digest=sha256:d9b56fc331aaa3120535c29f4e1105d849456da2e72b7bef643a33a9a7d0b059

Observation 376adc86-d204-4089-8f9e-cecd44d61290 · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale.

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders An image is worth 16x16 words: Transformers for image recognition at scale

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:37:33.752601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:37:17.620244Z digest=sha256:209f6cc18ac002917b35623c34a124b42c59185d35894a8a5cfe0837eab1f250

Observation 2676de86-cf63-4741-8c40-e4de5cae84b2 · outbound

This paper cites Vlmevalkit: An open-source toolkit for evaluating large multi- modality models.

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Vlmevalkit: An open-source toolkit for evaluating large multi- modality models

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:37:33.587582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:37:17.746820Z digest=sha256:df7dc28a2a6de5ab885d7e5f0ee986f48a5c0907e2b84dc7ace0a09f5b151b9f

Observation 7f21be05-08c7-49b4-9be9-780019e38bd2 · outbound

This paper cites Slowfast networks for video recognition.

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Slowfast networks for video recognition

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:37:33.431518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:37:17.902941Z digest=sha256:45d7d9325decf7587e9bcd4e8c3e3164bf62f03b4250f9374bcfd3008dfcb9f3

Observation 809b72ba-d3a4-41c5-8700-937a966fc819 · outbound

This paper cites Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis.

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:37:33.155237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:37:18.023177Z digest=sha256:10227dcfff92f0290e7e05580bf735119ca1703b68dee2afebfd9de2f60b6471

Observation 47dd4bc0-84f5-4db0-b975-1d6a06d65372 · outbound

This paper cites The Llama 3 Herd of Models.

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders The Llama 3 Herd of Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:18.121807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:18.121807Z digest=sha256:0d558775a5b65be4d8813cb987a46941932956ecdbdf596756a994206c4d2f2e

Observation 2381eef4-f583-4f76-94e7-c73d273512bf · outbound

This paper cites M-llm based video frame selection for efficient video understanding.

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders M-llm based video frame selection for efficient video understanding

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:37:32.959706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:37:18.286872Z digest=sha256:f8cbe670243991e4783d834fcc7a3c46bc6a3fc923011b2e8c8ca0982b32ea76

Observation 3750f7a5-fdfc-4ed9-aafa-7789b8c5546b · outbound

This paper cites Chat-univi: Unified visual representation empowers large language models with image and video understanding.

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Chat-univi: Unified visual representation empowers large language models with image and video understanding

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:37:32.786989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:37:18.409318Z digest=sha256:409366d826c9fd7afbe10f853e32c9ae9729149c85740ce500f36539ae03ba45

Observation fa668bc0-cee4-43e7-b805-46e964487e50 · outbound

This paper cites Language repository for long video understanding.

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Language repository for long video understanding

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:37:32.625784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:37:18.555490Z digest=sha256:7d8ebbeb39fb78e446b91e04c662de8ce87b853932e834c686cf11bc2a5c17ee

Observation a8d39fa3-b1cf-4f27-b2a2-2587ffad92a7 · outbound

This paper cites An image grid can be worth a video: Zero-shot video question answering using a vlm.

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders An image grid can be worth a video: Zero-shot video question answering using a vlm

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:37:32.441482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:37:18.682198Z digest=sha256:fe3777833f700116b24de61756b216f40360b0f085de5defeab4370a41818ccd

Observation 4241b8f0-2af6-46c9-9a51-1bf2b12aaab9 · outbound

This paper cites Lmms-eval: Accelerating the development of large multimoal models, March 2024.

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Lmms-eval: Accelerating the development of large multimoal models, March 2024

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:37:32.178385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:37:18.798283Z digest=sha256:81fbc03c639e9ab4fb87d96d2b0c03ac38ec05f8601064a90b31dcfac88a2068

Observation 291baed7-2303-4cf2-977a-841f9a3c3bf3 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders LLaVA-OneVision: Easy Visual Task Transfer

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:18.912546Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:18.912546Z digest=sha256:b30e85ed06f3c983a0538a07b9d765229a80211103d8c0e27a556aad32649e51

Observation 5d203498-19e8-4011-bb17-7995405bec65 · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:37:31.953262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:37:18.979088Z digest=sha256:cea543526889e8a4482eae718aadf32c3da9013b78e2d8934249414bf7d76a27

Observation 60b66c9c-95ff-4c99-aa4f-806c7132fb93 · outbound

This paper cites VideoChat: Chat-Centric Video Understanding.

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders VideoChat: Chat-Centric Video Understanding

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:19.138328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:19.138328Z digest=sha256:06e2c832399e9ca333ebfd1e4ce78827c8126f64e4c37528b57ef3462eedb4ab

Observation 01bfaaa7-1cec-48f4-8706-3dc35bce83d9 · outbound

This paper cites Llama-vid: An image is worth 2 tokens in large language models.

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Llama-vid: An image is worth 2 tokens in large language models

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:37:31.759563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:37:19.222339Z digest=sha256:7b99bd67c180f4c81492c75e1647c77daa29b8c13741dd215aebf7bd7017864e

Observation 5d5a836a-f36a-4515-98b8-f0983000e696 · outbound

This paper cites Video-llava: Learning united visual representation by alignment before projection.

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Video-llava: Learning united visual representation by alignment before projection

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:37:31.614389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:37:19.291562Z digest=sha256:5ae8ae602d9a098482b03883365bbd591e115a3a9468177bf89cb092bb396e24

Observation 638b84e3-ccde-4770-8b05-ab1b71fd4652 · outbound

This paper cites Vila: On pre- training for visual language models.

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Vila: On pre- training for visual language models

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:37:31.426866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:37:19.353735Z digest=sha256:ee88de6ace09704cb5bfe615778f5d24f27fa76c132f02e5f2ae7c14374adc5a

Observation c00adb3d-44ed-426f-a8e7-1acd59c1535a · outbound

This paper cites Llava-next: Improved reasoning, ocr, and world knowledge, January 2024.

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Llava-next: Improved reasoning, ocr, and world knowledge, January 2024

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:19.424939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:19.424939Z digest=sha256:be076bd9fdc68c89aeb49b3c7e27ea0da6a434a66466874facf9291ada2d889f

Observation e16f8ea2-1456-4418-80ed-0254c24804d7 · outbound

This paper cites Visual instruction tuning.Advances in neural information processing systems (NeurIPS), 36:34892–34916, 2023.

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Visual instruction tuning.Advances in neural information processing systems (NeurIPS), 36:34892–34916, 2023

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:37:31.189946Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:37:19.493444Z digest=sha256:bfb2a4654b646ec4d33981806b12d70152d90b2be0fe4e9802c25257b587e655

Observation 6954443f-d472-471a-ba74-4e194b2f62cf · outbound

This paper cites Timecraft: Navigate weakly-supervised temporal grounded video question answering via bi-directional reasoning.

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Timecraft: Navigate weakly-supervised temporal grounded video question answering via bi-directional reasoning

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:37:30.918499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:37:19.566948Z digest=sha256:098a3f68ccbc52290b9f62dad69984d11bb1b5cd1fc94c391fafccefc1b4b728

Observation b426fb8e-9c4a-44d3-9d07-e7078607e366 · outbound

This paper cites Lost in the middle: How language models use long contexts.

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Lost in the middle: How language models use long contexts

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:37:30.659000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:37:19.644740Z digest=sha256:3e06b38e9694bd8d15e9cdc67f347a8450c4d320cac28619d09a6c713738af27

Observation 2f4cc9f1-282b-41ed-9b30-625553421422 · outbound

This paper cites St-llm: Large language models are effective temporal learners.

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders St-llm: Large language models are effective temporal learners

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:37:30.370543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:37:19.721077Z digest=sha256:50bdd2c5e9b601435d4cfe3252bf35e51740e374568b963cf74b0ddd1ead6be3

Observation 81f55e11-6d72-4da8-a94c-73e27f7c50ec · outbound

This paper cites Bolt: Boost large vision-language model without training for long-form video understanding.

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Bolt: Boost large vision-language model without training for long-form video understanding

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:37:30.122315Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:37:19.791097Z digest=sha256:3d227c38985f4e17de2896d1b68f9b5b62067738559f05a85c3def2418051c3a

Observation 5342c0cd-7585-46c2-a601-c2740e0a62c8 · outbound

This paper cites Drvideo: Document retrieval based long video understanding.

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Drvideo: Document retrieval based long video understanding

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:37:29.987173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:37:19.859288Z digest=sha256:85dcf315edee4354eecd32501d109722d0c031a7b02693cbfa482c89cd1ca75b

Observation d4d72409-db44-44cc-b222-457dff76e831 · outbound

This paper cites Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models.

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:19.932480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:19.932480Z digest=sha256:293987a62d44918250f90b5d363dd37a661c32c4a4d326be49a9a3d461193ea8

Observation ffbf24f6-cd17-462d-8bb0-3c82bb673924 · outbound

This paper cites Egoschema: A diagnostic benchmark for very long-form video language understanding.

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Egoschema: A diagnostic benchmark for very long-form video language understanding

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:37:29.773619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:37:20.004759Z digest=sha256:7e8e06e3a00161435e83e29ccde74a8463ea10f395f491ff6a605dbb1d8780c3

Observation ae471a22-25b4-492d-a176-bca78cc195e3 · outbound

This paper cites Morevqa: Exploring modular reasoning models for video question answering.

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Morevqa: Exploring modular reasoning models for video question answering

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:37:29.610912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:37:20.072028Z digest=sha256:18b149bca8c982cae7c6e3974e9f14ff36ac893dcb93dc3b9ff89f51eed18643

Observation 0abc113c-9916-4985-825b-dc805945947b · outbound

This paper cites Branch-and-bound algorithms: A survey of recent advances in searching, branching, and pruning.

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Branch-and-bound algorithms: A survey of recent advances in searching, branching, and pruning

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:37:29.526632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:37:20.145930Z digest=sha256:c84ef145bfc16bd40b572fd740115044c820336c60482e67d8a30825d3a92beb

Observation d1c22e62-5148-4f3c-bd01-6c185a24dec6 · outbound

This paper cites Chatgpt: Optimizing language models for dialogue, 2023.

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Chatgpt: Optimizing language models for dialogue, 2023

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:37:29.389856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:37:20.236870Z digest=sha256:fc0d50d177076adbf1b21e5f3b8b9969f6e4cbd151eeb6d2426a5ff2efa62d23

Observation 02da2d3a-496d-4709-bcc1-4a52e6737496 · outbound

This paper cites Too many frames, not all useful: Efficient strategies for long-form video qa.

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Too many frames, not all useful: Efficient strategies for long-form video qa

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:37:29.235070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:37:20.321049Z digest=sha256:83147eb1975dca1aaa9213f42f55319414a7e11bb5e7f6e11417e21cd98c6940

Observation deb2fcb6-9589-4fa4-b275-b8777c51689d · outbound

This paper cites Momentor: Advancing video large language model with fine-grained temporal reasoning.

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Momentor: Advancing video large language model with fine-grained temporal reasoning

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:37:29.090406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:37:20.447778Z digest=sha256:83128960301b1383ce05a5dd8e66f120fda51b968c345b503b0e3e0c2a4ad44e

Observation 9347cb9c-82d3-4a9e-bf0e-53f2dcf2f7f3 · outbound

This paper cites Learning transferable visual models from natural language supervision.

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Learning transferable visual models from natural language supervision

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:37:28.954666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:37:20.542794Z digest=sha256:9901bf04b140b330e8d17cfd1d826bc38f890a53f553605dee6bdb8f24954919

Observation 3bc64dc1-212e-4d68-9cf6-0e9ba1c2a852 · outbound

This paper cites The knapsack problem: a survey.

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders The knapsack problem: a survey

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:20.635358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:20.635358Z digest=sha256:a21d57275f3058b715887593478ab5bd18913c6b15438d704beb440c2ba6ba09

Observation 312e0497-9293-483e-9372-348a673c1466 · outbound

This paper cites LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding.

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:20.727692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:20.727692Z digest=sha256:d0fd613ce232310b783c2e408274b4512366a03a1af0e896d4a993762d84e2b2

Observation f19b5cbc-5429-425a-8010-6d84e9d4ed61 · outbound

This paper cites Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding.

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:20.806587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:20.806587Z digest=sha256:c5a59fd9756d54b4e1cb0150afafd0b507d6426e799b56d61ed87409f38891c8

Observation fbd4598c-ac5d-4e0c-aeba-4bd4cc9806cf · outbound

This paper cites Two-stream convolutional networks for action recognition in videos.

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Two-stream convolutional networks for action recognition in videos

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:37:28.843270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:37:20.873682Z digest=sha256:4ba9c70edc0875b5f9e689f521639888a37e39b36fd787f0fd454a60ca3757e4

Observation d462dcf3-d607-4bff-9bed-dd3901c0c289 · outbound

This paper cites Moviechat: From dense token to sparse memory for long video understanding.

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Moviechat: From dense token to sparse memory for long video understanding

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:37:28.689205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:37:20.984827Z digest=sha256:1660fb49b455be8fd9fb4f27dd14233095f28100fb43d3ec510f324a96715cce

Observation 7b755651-2206-48b2-9d89-86b81dba6344 · outbound

This paper cites Mdp3: A training-free approach for list-wise frame selection in video-llms.

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Mdp3: A training-free approach for list-wise frame selection in video-llms

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:21.085807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:21.085807Z digest=sha256:44f43c339d4d87207a6a103a4d5f71916291a2fe5123afeef2490772e1507958

Observation 542d090a-bde5-476e-93c8-631aaba47c51 · outbound

This paper cites Adaptive keyframe sampling for long video understanding.

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Adaptive keyframe sampling for long video understanding

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:37:28.544301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:37:21.159282Z digest=sha256:dacef6140c34f4fb5ef58e26c758e1e1cc12f26a2c2ca907b7ed0281f878c339

Observation fb5be959-6296-4046-9e8b-69de40a997a2 · outbound

This paper cites Gemma: Open Models Based on Gemini Research and Technology.

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Gemma: Open Models Based on Gemini Research and Technology

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:21.246351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:21.246351Z digest=sha256:15fac6a878b26462684c7b33a7985852f65cddfb87c9d65ec7a2ee3f009e897f

Observation 0bb7053e-8653-4f74-a8ed-f522d94005a6 · outbound

This paper cites Cambrian-1: A fully open, vision- centric exploration of multimodal llms.

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Cambrian-1: A fully open, vision- centric exploration of multimodal llms

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:37:28.389747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:37:21.360353Z digest=sha256:6e132885c05b98fa6108dd48511ab21a53b707cf9c12614fecb5e5e9d96891f8

Observation a09a5b92-a3c3-46d7-9c28-29122d677e3c · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:21.439956Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:21.439956Z digest=sha256:7cc36d2db7c5d439895c5c5043c54acaf6b26a817594b550da13b4804e2bcda9

Observation 46bcfb17-ae86-4fea-9e96-e026cddb3bfc · outbound

This paper cites Attention is all you need.

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Attention is all you need

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:37:28.221238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:37:21.559745Z digest=sha256:d64f2679b987791a963e7849cb224519e7be8d6a5335c921119f5aea647e9228

Observation a9736535-42eb-4e5f-8817-6603449fa2e5 · outbound

This paper cites Show and tell: A neural image caption generator.

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Show and tell: A neural image caption generator

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:37:28.070943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:37:21.692309Z digest=sha256:73e46086710142a73d62bb42c28c5bb3603428a9203eba68c9efe4b970aae9f0

Observation cf88f0d5-6501-4dae-b17b-c9deec7298aa · outbound

This paper cites Efficient large language models: A survey.Transactions on Machine Learning Research (TMLR), 2024.

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Efficient large language models: A survey.Transactions on Machine Learning Research (TMLR), 2024

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:37:27.896903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:37:21.768128Z digest=sha256:7a540948dbf9cd99c44cb9eefa8c2af928763f42777f62c3d7bbc00793408e95

Observation 8e98f985-ca8e-4fc2-a759-67f3c756a9be · outbound

This paper cites Weakly supervised gaussian contrastive grounding with large multimodal models for video question answering.

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Weakly supervised gaussian contrastive grounding with large multimodal models for video question answering

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:37:27.715659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:37:21.846512Z digest=sha256:43fbb91bd558621e7e546cc49a20af201efa7923eb4d34a58b9d14b527f24d03

Observation 88731b3a-d85e-4308-ab9d-076a50bdcc2a · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:21.942911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:21.942911Z digest=sha256:40eee836f4817bf082d266d315a322244255c83170e24f61212a621f3dcaed34

Observation 26926a99-6fe9-4763-ae56-44f168a64c7d · outbound

This paper cites ReTaKe: Reducing Temporal and Knowledge Redundancy for Long Video Understanding.

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders ReTaKe: Reducing Temporal and Knowledge Redundancy for Long Video Understanding

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:22.035335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:22.035335Z digest=sha256:33dc5ce8945cb8cc2c588722c7b0b296492d23a7914932f531867382ca68cd61

Observation ae110e84-5baf-4ba8-94c9-2236d33e803d · outbound

This paper cites Videoagent: Long-form video under- standing with large language model as agent.

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Videoagent: Long-form video under- standing with large language model as agent

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:37:27.552815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:37:22.139175Z digest=sha256:79d106824e3618e94f7717faf5c71897a1b760c0bb294c40ef1b6ed3ed56c244

Observation eb9742ac-5b10-4d64-a3ed-357f9bcc1d79 · outbound

This paper cites Videotree: Adaptive tree-based video representation for llm reasoning on long videos.

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Videotree: Adaptive tree-based video representation for llm reasoning on long videos

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:37:27.342292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:37:22.236365Z digest=sha256:dfda75df33ba235c6d1e3fbefefda0d4f845de502ca09b3bc2c55c6b72e97e8c

Observation 7cd90885-7735-4b77-8764-6ba6759e772e · outbound

This paper cites Dibs: Enhancing dense video captioning with unlabeled videos via pseudo boundary enrichment and online refinement.

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Dibs: Enhancing dense video captioning with unlabeled videos via pseudo boundary enrichment and online refinement

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:37:27.171688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:37:22.338769Z digest=sha256:6c7927f8ea3172c38e8620f7547c4d740b534a24ec8700c372c4ffd474f95dd5

Observation 061e4745-6e27-4634-a3ca-076def2997d1 · outbound

This paper cites Longvideobench: A benchmark for long-context interleaved video-language understanding.

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Longvideobench: A benchmark for long-context interleaved video-language understanding

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:37:26.974968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:37:22.450389Z digest=sha256:1e8ddc599eea3512c77b978854d2b0a08133babec9abc006a81c7270ae3bedfd

Observation 7998ee8f-53e2-406a-944f-7d9d6e446f8b · outbound

This paper cites Next-qa: Next phase of question-answering to explaining temporal actions.

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Next-qa: Next phase of question-answering to explaining temporal actions

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:37:26.807231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:37:22.581586Z digest=sha256:9b46144370f3294bc2bb6abc28f012b273781c3006dfc75ecc9fc78e2249bafc

Observation 519ac6c8-ad9d-415d-bb49-44f01783acce · outbound

This paper cites Can i trust your answer? visually grounded video question answering.

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Can i trust your answer? visually grounded video question answering

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:37:26.664010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:37:22.691030Z digest=sha256:989a75ae5accf28ed66fe218692b5c915dc364e80ca0424da0216e03a3de6c73

Observation c3b083f2-ee51-4cb1-a99d-b3b5c2ef56c0 · outbound

This paper cites Effective long-context scaling of foundation models.

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Effective long-context scaling of foundation models

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:37:26.500839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:37:22.786503Z digest=sha256:9205d52cff67dc8904c1f5027230b7964a821758344330ac5d0bd6001e96a853

Observation a990e067-485f-4cf9-852f-ee77af800c35 · outbound

This paper cites PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning.

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:22.900514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:22.900514Z digest=sha256:f5511e3df2499e3b8efd6d31bdf5e68329b50c0708e874c8845dc677c9deaf7d

Observation cba55f33-db7f-408d-99fa-78317f5baaca · outbound

This paper cites SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models.

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:23.032653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:23.032653Z digest=sha256:c7605944db25f0b3227cd57a3acc56e146f84d3144f1bd35707b9e8f2d8dd956

Observation cb3d9dd3-06aa-45e8-bdb6-fdb75c4eda2e · outbound

This paper cites Zero-shot video question answering via frozen bidirectional language models.

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Zero-shot video question answering via frozen bidirectional language models

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:37:26.332155Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:37:23.130555Z digest=sha256:c814fcec9a9adb3bcc0d425f99c85a266f624ba41904891cfb407c2b3c6738ee

Observation 51652350-edb5-4a59-8448-5a2ab454b129 · outbound

This paper cites Vid2seq: Large-scale pretraining of a visual language model for dense video captioning.

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Vid2seq: Large-scale pretraining of a visual language model for dense video captioning

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:37:26.186839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:37:23.230304Z digest=sha256:b1f2892bd9586acc05d94fdbf736bde6100b74a6fa58b2a656919393b1d7cc93

Observation 9298836e-13e0-461f-ab6c-1abfb7ea459e · outbound

This paper cites Dense connector for mllms.

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Dense connector for mllms

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:37:26.019452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:37:23.342782Z digest=sha256:5dc57619606c9316876d4d218dc0a0e623bd0eb1351b06ae2142af41539d0748

Observation 0ef5db10-d2dc-4306-898f-c0a74ce8d362 · outbound

This paper cites Generative Frame Sampler for Long Video Understanding.

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Generative Frame Sampler for Long Video Understanding

Reference 73

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:37:24.889986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:37:23.440876Z digest=sha256:2ebc08a230c01a4d0abc0429034f297e8cd99f6e47a05abca9b8aaa4fcf47ad2

Observation e9d53f51-3628-4ccd-8161-bafbb69f4d9a · outbound

This paper cites MiniCPM-V: A GPT-4V Level MLLM on Your Phone.

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:23.584263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:23.584263Z digest=sha256:1847d71ae2c51063ebea9f620b1b4da7f44c97b961d76dfc1698272fb1656228

Observation 1388fada-bbd9-4c81-83ce-e5b474ed503f · outbound

This paper cites Self-chained image-language model for video localization and question answering.

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Self-chained image-language model for video localization and question answering

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:37:25.913672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:37:23.714914Z digest=sha256:79cd388eb827573e4aacac52d8307686ae1f67954b2025b625cbfb5d335ba43b

Observation a2d295cf-2c3e-4d3b-864f-226d08d397da · outbound

This paper cites Frame-voyager: Learning to query frames for video large language models.

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Frame-voyager: Learning to query frames for video large language models

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:37:25.779824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:37:23.826563Z digest=sha256:767e9b7aee4261a56c1f9f8e9d99bdea101e4943111d92b7e98571b8a4267644

Observation 49e98b79-f970-452f-9e47-f1468db5e4da · outbound

This paper cites Sigmoid loss for language image pre-training.

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Sigmoid loss for language image pre-training

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:37:25.554077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:37:23.882899Z digest=sha256:36277892b53c94c105791a7604d2bd5ac934206cab3040633d2d54647912adc1

Observation 842a74e2-9b22-453c-b188-65da3f3cc9e1 · outbound

This paper cites A simple llm framework for long-range video question-answering.

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders A simple llm framework for long-range video question-answering

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:37:25.380735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:37:23.986712Z digest=sha256:b0a9c8034d283c71c8c0e99d9826581f8f6bb5d280c62d20c35170810d52d4c9

Observation d0f16f33-24b5-4af9-ad4e-4b8e07355c6d · outbound

This paper cites Long Context Transfer from Language to Vision.

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Long Context Transfer from Language to Vision

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:24.093163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:24.093163Z digest=sha256:d5718d7008eece9bdc21b38a5e6f1164a0f56b8f808b2d79ecba2a45fa83065c

Observation 79b9db04-5fc5-48b4-aadf-697adeb8f07e · outbound

This paper cites Llava-next: A strong zero-shot video understanding model, April 2024.

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Llava-next: A strong zero-shot video understanding model, April 2024

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:24.190869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:24.190869Z digest=sha256:86f3907be081982bc6e08a55a4aa89b672f6d4260113d89674dded6b8146a5f1

Observation b30ebedc-b77a-4463-b3a6-ef8c7a773f6d · outbound

This paper cites LLaVA-Video: Video Instruction Tuning With Synthetic Data.

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:24.258273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:24.258273Z digest=sha256:f1f121e490be05652eecef55116547a19758369b30692c379d4bf9a2d9b08264

Observation b028f069-8365-4ff5-903f-0d9928d4f313 · outbound

This paper cites MLVU: Benchmarking Multi-task Long Video Understanding.

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders MLVU: Benchmarking Multi-task Long Video Understanding

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:24.363162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:24.363162Z digest=sha256:a504f19dd4d90ec8482453383866d90b00d797ad143ec94bd3fa4502f666edfb

Observation 4793def1-2e8a-4082-a776-2a2c5df55d59 · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:24.461768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:24.461768Z digest=sha256:aa2cc8c4bede20b802153777ad3c5ac53122f458781f4b716f031639aeb59f61

Observation 447bbede-b0be-4fbc-9feb-22827ba5a7bd · outbound

This paper cites Apollo: An Exploration of Video Understanding in Large Multimodal Models.

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Apollo: An Exploration of Video Understanding in Large Multimodal Models

Reference 84

Resolution
malformed identifier
no resolver link, observed 2026-08-07T12:37:24.588694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:24.588694Z digest=sha256:eead4469e06d77c869993e4971c5e2e7d317001f612e687aefae3886b05fb804

Pith citing papers

Observation 76dbb3a0-84e8-434c-9c6f-48b3029920af · inbound

Training-free Uncertainty Guidance for Complex Visual Tasks with MLLMs cites this paper.

Training-free Uncertainty Guidance for Complex Visual Tasks with MLLMs Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-04T13:27:56.359929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:27:56.359929Z digest=sha256:c56dc8a3d58a124ac37ea7d835979ace758c78076e031a707af055607d4370ca

Observation 767c7c7e-b77e-4aeb-8c50-e2dfe77869ac · inbound

CREST: Curvature-Regulated Event-Centric Sampling for Efficient Long-Video Understanding cites this paper.

CREST: Curvature-Regulated Event-Centric Sampling for Efficient Long-Video Understanding Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-12T03:06:18.772638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-12T03:06:09.753634Z digest=sha256:bced75dddb6c92c02424c8e7e743f68a12edd392f640b530f3d47235acd3d0b4

Observation df5b6b56-6f9f-4886-9414-484068385707 · inbound

One Ranking, Any Budget: Matryoshka Evidence-to-Context Frame Selection for Long-Video Understanding cites this paper.

One Ranking, Any Budget: Matryoshka Evidence-to-Context Frame Selection for Long-Video Understanding Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T14:39:08.699220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:39:08.699220Z digest=sha256:7fe18c226ec8be99aab02e07e448d306698343b7f2b05d6f1b621d63c6bb3098