Pith. sign in

Paper Citation Record · LEDGER

Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs

As of 20 August 2026, this Paper Citation Record lists 42 of 42 outbound references and 9 inbound Pith citation observations for arXiv:2506.22139.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.22139 v3

Coverage vector

measured 42 of 42 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T22:15:05.248695Z

measured 51 of 51 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T23:42:35.137666Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T16:09:57.149192Z

Reference resolution

42 of 42 outbound references displayed

  • verified exact0
  • verified fuzzy21
  • unresolved21
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a2c548b6-25d1-44d2-ac91-7ec88012bc3a · outbound

This paper cites Qwen Technical Report.

Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs Qwen Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T22:15:01.915030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:15:01.915030Z digest=sha256:1ceac513fc6f59c21d59065fd3f1f2301958e76273c8fab415081545c263711c

Observation ea63bb34-9624-424a-a837-d4b9cafa0642 · outbound

This paper cites Robust motion-guided frame sam- pler with interpretive evaluation for video action recognition.

Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs Robust motion-guided frame sam- pler with interpretive evaluation for video action recognition

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:15:09.053397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T22:15:01.967902Z digest=sha256:7083434fe01a503b23147d4c4007fe3cfffda9d9d9c2c3a91e0278ba06d38a50

Observation 86249126-8260-4bac-8d4e-d5b7523ba862 · outbound

This paper cites Sharegpt4video: Improving video understanding and generation with better captions.

Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs Sharegpt4video: Improving video understanding and generation with better captions

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:15:08.884299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T22:15:02.033532Z digest=sha256:090944eed4e6b708999c2cfd979f384e06776b5e4c9a5d4225e021268547e1aa

Observation 5ca6dea2-4ce0-4342-9bda-680746d4e6d6 · outbound

This paper cites Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks.

Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T22:15:02.128356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:15:02.128356Z digest=sha256:6938482f72594a9d97a78e26fbf8f0f9c582c41096623196e14747c416242151

Observation 636f60ca-4311-4df5-bf28-12d428f28cbd · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T22:15:02.241212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:15:02.241212Z digest=sha256:5821d59b2bdfbbcd28a48477d8d00708050ec659532f37aaf9227c7c03e619ba

Observation 55a5ff53-44d3-4273-8475-82950400326c · outbound

This paper cites Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos.

Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T22:15:02.340057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:15:02.340057Z digest=sha256:37dafca7eaa4bc8a91f2e04e91e86a068e10b965128bfad960633c0281515ef7

Observation 260f516a-eb60-4a70-aa94-5a542647b211 · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T22:15:02.450017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:15:02.450017Z digest=sha256:a8793447e68413f1e7e6b878ffdf99d5222b0bd71e931e110716680390532379

Observation d7ed91ed-796f-4940-b039-802d573b9f48 · outbound

This paper cites Categorical Reparameterization with Gumbel-Softmax.

Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs Categorical Reparameterization with Gumbel-Softmax

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T22:15:02.527447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:15:02.527447Z digest=sha256:7ccfa4b9e468f15e2a4e7f5bf3c1e3a9cb0d3e1f5719dcceb6fb9d0735b953a8

Observation e1403065-aac9-4fcd-bfdc-403074e446ef · outbound

This paper cites Chat-univi: Unified visual representation em- powers large language models with image and video un- derstanding.

Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs Chat-univi: Unified visual representation em- powers large language models with image and video un- derstanding

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:15:08.677569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T22:15:02.615899Z digest=sha256:0001f4e9745ced579b69d9cd6ba31ffc22991be79ddab2b1487aafe31c1295ea

Observation 49eacf54-21f3-4a01-b6c3-b164f5b11144 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs LLaVA-OneVision: Easy Visual Task Transfer

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T22:15:02.714570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:15:02.714570Z digest=sha256:bf7d6cadbf5aace96f0b6c5510043ec830d8b4ae43cfd0aee82c872fd846c23d

Observation f53e097e-80e5-4101-b645-9a7a49a929b7 · outbound

This paper cites Mvbench: A comprehensive multi-modal video understand- ing benchmark.

Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs Mvbench: A comprehensive multi-modal video understand- ing benchmark

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:15:08.537709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T22:15:02.778441Z digest=sha256:0b40bf9f11c354e4a313be512dceca0268a2855281a88a15812f5b018e403f7f

Observation a797c863-8993-4d59-b303-c2213ace2331 · outbound

This paper cites Mon- key: Image resolution and text label are important things for large multi-modal models.

Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs Mon- key: Image resolution and text label are important things for large multi-modal models

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:15:08.389761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T22:15:02.859142Z digest=sha256:443b1706567867701e9797e4eec8aa5ce72ea603a99d54c7f9cdea7c4330514c

Observation 76dc95d8-f7e4-4b88-9a02-94e10a7c9b77 · outbound

This paper cites KeyVideoLLM: Towards Large-scale Video Keyframe Selection.

Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs KeyVideoLLM: Towards Large-scale Video Keyframe Selection

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T22:15:02.922173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:15:02.922173Z digest=sha256:2618ca30ea4360ef744ce5362920ae43d564d4bd1414726fe3c22a75d4318fa2

Observation b2dd8b82-c552-47d0-885a-c4f771f400d5 · outbound

This paper cites Video-llava: Learning united visual representation by alignment before projection.

Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs Video-llava: Learning united visual representation by alignment before projection

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:15:08.243146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T22:15:02.991356Z digest=sha256:9314c81a67ee75881a6768b8878eea2b4d49001f1f8d3f8986dc91f127ea12f2

Observation 2475e21a-6cd2-4261-8e08-9faca5ee815c · outbound

This paper cites Vila: On pre-training for vi- sual language models.

Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs Vila: On pre-training for vi- sual language models

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:15:08.048504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T22:15:03.062699Z digest=sha256:091de25179b7789c39bf2a0c226ad20ff5112323d28549c0dd5359be004c9d6f

Observation d5d2932f-6139-47ee-8d33-e1b0201e24e6 · outbound

This paper cites Visual instruction tuning.

Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs Visual instruction tuning

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:15:07.838909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T22:15:03.144052Z digest=sha256:e4091a6b1b22e5f6218073c747621eff03c407ba84226c7e36951049f328648f

Observation 73badf97-5289-4ab2-9a12-ae2e30d2d29d · outbound

This paper cites Improved baselines with visual instruction tuning.

Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs Improved baselines with visual instruction tuning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T22:15:03.202524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:15:03.202524Z digest=sha256:0d16f154d1fabe0a7742f9de9eac5ac77a03ec7d426711819432b08dd5dc7f8d

Observation b8379fdf-f320-447c-9623-29220e3249d2 · outbound

This paper cites Llavanext: Improved reasoning, ocr, and world knowledge, 2024.

Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs Llavanext: Improved reasoning, ocr, and world knowledge, 2024

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:15:07.653968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T22:15:03.286726Z digest=sha256:bf703861ddd23ec86321c6a393df4ea4ffc9a2762563607aa6d2bf77fe20fef5

Observation 0f33d806-78f9-4ccf-9bd3-34caa259491c · outbound

This paper cites Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input.

Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T22:15:03.351183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:15:03.351183Z digest=sha256:f995922c4abea91be21c597e91089283c898b1613a0c3c602e81cb619ffee7f4

Observation 666db4a5-e386-4eac-b8a7-5681a16d0392 · outbound

This paper cites Video-chatgpt: Towards detailed video understanding via large vision and language models.

Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs Video-chatgpt: Towards detailed video understanding via large vision and language models

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:15:07.490542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T22:15:03.417538Z digest=sha256:3dc2a795955f084147799f5cceebd068a2f8cb2c3ee64d37ca3223007eef4fbf

Observation ad62d436-89cf-4000-b98e-3e65e6ed29ca · outbound

This paper cites Hello gpt-4o, 2024.

Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs Hello gpt-4o, 2024

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:15:07.287773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T22:15:03.481220Z digest=sha256:b7fdabaef488dca66ff65a8487b58e9e0407735d1689fe175ecab86d45ceb924

Observation 9fbe48c4-ba7c-415b-a164-a6f7af190310 · outbound

This paper cites Learn- ing transferable visual models from natural language super- vision.

Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs Learn- ing transferable visual models from natural language super- vision

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:15:07.110497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T22:15:03.561078Z digest=sha256:d4232b16397246212e4a91f4f7cf1579f9d61da649d9fed63022f4bfafeb102d

Observation 20018ddc-e8da-4d12-8dd6-b41804ca24cd · outbound

This paper cites Timechat: A time-sensitive multimodal large lan- guage model for long video understanding.

Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs Timechat: A time-sensitive multimodal large lan- guage model for long video understanding

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:15:06.906491Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T22:15:03.678869Z digest=sha256:eb961566fb9a06b1b0cdf1bc06a3b79413a894d3e5e2ff55cb2c2d89141c62da

Observation 51b1b8f9-be18-4d35-bc41-a649916768d0 · outbound

This paper cites Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding.

Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T22:15:03.755141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:15:03.755141Z digest=sha256:41ba776680edd34e5181101d066a4a7af07d16a931f1274a8cab27e6cb27b0c4

Observation 5f3b6777-e92f-4325-91a5-bf8b25fc6660 · outbound

This paper cites Moviechat: From dense token to sparse memory for long video understanding.

Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs Moviechat: From dense token to sparse memory for long video understanding

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T22:15:03.824422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:15:03.824422Z digest=sha256:ebd8ee07f8b1fd0e8c002d68bc4e299435c65f9bd81f420a21592ee8abaa2f3d

Observation b47ab2da-1955-449d-998b-88df156c4469 · outbound

This paper cites Video understanding with large language models: A survey.

Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs Video understanding with large language models: A survey

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T22:15:03.875000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:15:03.875000Z digest=sha256:7106cc567739628896c357500b7382f0c592e16dd6301153ca027b4bb107becb

Observation e4359f59-bc8b-46fe-a718-628a619f961b · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T22:15:03.971684Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:15:03.971684Z digest=sha256:ff77fdc67bc352a3bb8096ebb14c4a16bbbd9ccfe37bae09ccbca8657a4b92d7

Observation 1f4c3484-0531-472f-ad6f-3aa4f348d22f · outbound

This paper cites Internvideo2: Scaling foundation models for multi- modal video understanding.

Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs Internvideo2: Scaling foundation models for multi- modal video understanding

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:15:06.755357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T22:15:04.029772Z digest=sha256:6934c9daceee9dc55170fdce70035bada77fc9a961b0958d649df0a47090ab70

Observation e39c4189-926f-46bb-b906-7f1542fec27a · outbound

This paper cites Longvideobench: A benchmark for long-context interleaved video-language understanding.

Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs Longvideobench: A benchmark for long-context interleaved video-language understanding

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:15:06.629724Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T22:15:04.102916Z digest=sha256:95883271fecc69402549780e90c9f42378836f231271584762f433ab1321723b

Observation e73ebea1-c1cf-4f09-8858-be514963c3a3 · outbound

This paper cites Multi-agent reinforcement learning based frame sampling for effective untrimmed video recognition.

Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs Multi-agent reinforcement learning based frame sampling for effective untrimmed video recognition

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:15:06.502690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T22:15:04.241569Z digest=sha256:98ea0a8ef1f4e40c00816880f6e410703b255dcc47e0ba3862450b314c5a419e

Observation d916726f-0327-4842-ae41-4fcc2571c33c · outbound

This paper cites Adaframe: Adaptive frame selection for fast video recognition.

Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs Adaframe: Adaptive frame selection for fast video recognition

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:15:06.387469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T22:15:04.299593Z digest=sha256:a19aef5c3452c22026f2218df083b1ea6d08c9eee8eb838e0c1f7b433716ef3c

Observation a4b2f519-101d-44cc-af46-640da2b1c2df · outbound

This paper cites LongVILA: Scaling Long-Context Visual Language Models for Long Videos.

Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs LongVILA: Scaling Long-Context Visual Language Models for Long Videos

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T22:15:04.382683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:15:04.382683Z digest=sha256:2cf0072ecbe31b69c86e3abe663ecb8d4449926e2d197aff70488c27e33c9736

Observation 665bd746-852e-4a8b-a786-607ecfbe37a9 · outbound

This paper cites Frame-voyager: Learning to query frames for video large language models.

Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs Frame-voyager: Learning to query frames for video large language models

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:15:06.261410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T22:15:04.480606Z digest=sha256:d437ec8d87ad9014bd16d4cefe7a060eb7a94890ee2a1d23cf1c4c8f4578ab36

Observation eb4d00fa-9cc6-44eb-bf77-de94c0c3682b · outbound

This paper cites Sigmoid loss for language image pre-training.

Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs Sigmoid loss for language image pre-training

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:15:06.143689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T22:15:04.561512Z digest=sha256:599cf8955e7230a4e4b524a4b8df3060dbe5a3e90fcdd795aad99a9301f89422

Observation e0e680ee-9a99-4961-857a-fd21ac5307dd · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T22:15:04.623326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:15:04.623326Z digest=sha256:1c4cbd7537f7e2e0d04b67a91b95f99280e72480fc9b258f74abb47dccd3fde2

Observation 5d985fbf-4b73-481f-b255-2b737e68829e · outbound

This paper cites LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models.

Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T22:15:04.686714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:15:04.686714Z digest=sha256:233d97c3b376ecc3a2f84a00fe44f3082935bd70b58f1eed1f1237ca6c6c2316

Observation ef8de3d9-0280-4974-bb2f-47853121def6 · outbound

This paper cites Long Context Transfer from Language to Vision.

Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs Long Context Transfer from Language to Vision

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T22:15:04.749344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:15:04.749344Z digest=sha256:ec836e3212fce8b6edd44401500caef68da6ab1af3434ea40118677985c77c4d

Observation 5a1f0d10-f715-47c1-b335-f47ffddde9c3 · outbound

This paper cites Direct Preference Optimization of Video Large Multimodal Models from Language Model Reward.

Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs Direct Preference Optimization of Video Large Multimodal Models from Language Model Reward

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T22:15:04.864133Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:15:04.864133Z digest=sha256:bf6fd151e0b9e87ad8e48b328090446914cafffbae87852d88b412b399701b91

Observation 0c307f2e-7cdd-4c08-b992-c4fc22cea03a · outbound

This paper cites LLaVA-Video: Video Instruction Tuning With Synthetic Data.

Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T22:15:05.019965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:15:05.019965Z digest=sha256:84372337a31b2b92d3be587e5c81c9e8514a7453f30f0dcd7699e1b7373c0a23

Observation 9fc18216-6abf-4dc7-915d-f115ba96b59f · outbound

This paper cites Mgsampler: An explainable sampling strategy for video ac- tion recognition.

Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs Mgsampler: An explainable sampling strategy for video ac- tion recognition

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:15:05.983536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T22:15:05.086874Z digest=sha256:ecafe09e5c19349799e6b92c18d7086ec74e7d44f7dd365716ab4119151b3af2

Observation f71645bf-2cfe-46ad-8924-4ae73dd3d962 · outbound

This paper cites MLVU: Benchmarking Multi-task Long Video Understanding.

Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs MLVU: Benchmarking Multi-task Long Video Understanding

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T22:15:05.172214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:15:05.172214Z digest=sha256:0928193a2af2dcbf0e6ce8257f9a675bbcdefab481258c3e2ce4c0a9652015ba

Observation 2eef322c-e13d-4329-90c3-8916f65c57a9 · outbound

This paper cites Limitations Q-Frame enhances query-aware video understanding, but it depends on pre-trained models, lacks explicit temporal modeling, and operates within a fixed token budget.

Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs Limitations Q-Frame enhances query-aware video understanding, but it depends on pre-trained models, lacks explicit temporal modeling, and operates within a fixed token budget

Reference 2024

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:15:05.867099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T22:15:05.248695Z digest=sha256:588e47e13736d0a1d76df9dd83ae62f9eca86f35cd1c2ed6e64ef8b5b7c5daf7

Pith citing papers

Observation 0b6ad907-50e4-4dc3-bfa1-ecd47740152c · inbound

AdaSpark: Adaptive Sparsity for Efficient Long-Video Understanding cites this paper.

AdaSpark: Adaptive Sparsity for Efficient Long-Video Understanding Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:56:04.269758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-10T17:21:47.439019Z digest=sha256:3cff8cf9a158221ca5acc252f218d52f4e5fb58f31150efea512f7e805b07421

Observation 33844b79-8431-4eda-bdce-8ca1e28b38dc · inbound

Multi-Scale Separable Fourier Neural Networks for Solving High-Frequency PDEs cites this paper.

Multi-Scale Separable Fourier Neural Networks for Solving High-Frequency PDEs Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs

Reference 38

Resolution
unresolved
no resolver link, observed 2026-07-12T15:34:37.002001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T15:34:37.002001Z digest=sha256:1714bd4999150bb58575a57cfbaa138d7e7432b7f9a442fd04cdaa9065c90ad0

Observation eacdb4a2-0578-4ee1-87d6-3cc48c57148e · inbound

Multi-Scale Separable Fourier Neural Networks for Solving High-Frequency PDEs cites this paper.

Multi-Scale Separable Fourier Neural Networks for Solving High-Frequency PDEs Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs

Reference 38

Resolution
unresolved
no resolver link, observed 2026-07-14T18:39:24.915547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T18:39:24.915547Z digest=sha256:712df104b0aa39d91070aa2f52daf3ff67b65c5be25f2c1f0f3c79a77bdcacad

Observation 079f6c34-65d1-42bd-aab7-82261c6a427e · inbound

PEEK: Picking Essential frames via Efficient Knowledge distillation cites this paper.

PEEK: Picking Essential frames via Efficient Knowledge distillation Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-07-01T19:16:01.318490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-28T22:46:53.704892Z digest=sha256:b6e42cc66df1d0357a2e1f70dd580fe09baa5c49b328ddf5c1bbbcdf46580678

Observation 9498d452-63c8-44f2-8693-0f8a72695bb3 · inbound

Zero-Shot 3D Question Answering via Hierarchical View-to-Token Transportation cites this paper.

Zero-Shot 3D Question Answering via Hierarchical View-to-Token Transportation Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs

Reference 26

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T02:26:27.027565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-06-28T10:54:02.188634Z digest=sha256:dfb419ed7fb988e06cc4d6d1642e9d4c25ba3581d3d5a181aa1035ef295fb343

Observation de4ae604-e237-4b7a-9565-2c18234cc729 · inbound

Watch, Remember, Reason: Human-View Video Understanding with MLLMs cites this paper.

Watch, Remember, Reason: Human-View Video Understanding with MLLMs Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:27:15.057718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-27T22:00:28.350003Z digest=sha256:a0198d4bab00b312ff3ff671d34305e65035f1d7aa13ca5f8839d107dcfce116

Observation 9824f3dc-f80b-4c09-b06c-904f35ade8c6 · inbound

Towards Fast and Effective Long Video Understanding of Multimodal Large Language Models via Adaptive Quasi-Gaussian Sampling cites this paper.

Towards Fast and Effective Long Video Understanding of Multimodal Large Language Models via Adaptive Quasi-Gaussian Sampling Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-07-04T16:09:57.150830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-26T00:53:43.629684Z digest=sha256:59cf3e80b34f7b5497eabbded08bd08bc7370bb5f8b60cf5821d5c69642e5782

Observation 5051ba87-dc57-4d7a-946d-cbf469c36426 · inbound

QCA: Query- and Content-Aware Keyframe Selection for Long Video Understanding cites this paper.

QCA: Query- and Content-Aware Keyframe Selection for Long Video Understanding Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs

Reference 38

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T14:17:02.668832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-02T14:09:54.549499Z digest=sha256:db6428897c0f816b0e11878aa0551333ba18361a9bbff4e697d71aa65714f381

Observation 83ce1718-bef9-4d79-a24a-13456167b10f · inbound

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding cites this paper.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.137666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.137666Z digest=sha256:5932bc5d49cd0e4b3ca748094c13351df151b54b147763e0a1bb23102e795204