Pith. sign in

Paper Citation Record · LEDGER

Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs

As of 8 August 2026, this Paper Citation Record lists 42 of 42 outbound references and 9 inbound Pith citation observations for arXiv:2506.22139.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.22139 v3

Coverage vector

measured 42 of 42 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T22:15:05.248695Z

measured 51 of 51 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T23:42:35.137666Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T16:09:57.149192Z

Reference resolution

42 of 42 outbound references displayed

  • verified exact0
  • verified fuzzy21
  • unresolved21
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a2c548b6-25d1-44d2-ac91-7ec88012bc3a · outbound

This paper cites Qwen Technical Report.

Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs Qwen Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T22:15:01.915030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:15:01.915030Z digest=sha256:2e3731f20d755b462257846eae318137133419ef036d29068bc25feaa4ec3343

Observation ea63bb34-9624-424a-a837-d4b9cafa0642 · outbound

This paper cites Robust motion-guided frame sam- pler with interpretive evaluation for video action recognition.

Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs Robust motion-guided frame sam- pler with interpretive evaluation for video action recognition

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:15:09.053397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:15:01.967902Z digest=sha256:3f5c4ffb67af4d0717f32a2b8702378382e755d623404b744dbf1e708206de94

Observation 86249126-8260-4bac-8d4e-d5b7523ba862 · outbound

This paper cites Sharegpt4video: Improving video understanding and generation with better captions.

Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs Sharegpt4video: Improving video understanding and generation with better captions

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:15:08.884299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:15:02.033532Z digest=sha256:d386133d5600d352cc153860fcc87da48b643e1fc870b4958075bd013a05c180

Observation 5ca6dea2-4ce0-4342-9bda-680746d4e6d6 · outbound

This paper cites Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks.

Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T22:15:02.128356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:15:02.128356Z digest=sha256:e37f3194842c5b665c9d7b856ff2cc5b77256add8f4093f2de2157ceb0f09847

Observation 636f60ca-4311-4df5-bf28-12d428f28cbd · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T22:15:02.241212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:15:02.241212Z digest=sha256:43cabe4324596bae024cb65c3dd315a5eaf32cdcff3062c7dfe6f7386d3b9685

Observation 55a5ff53-44d3-4273-8475-82950400326c · outbound

This paper cites Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos.

Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T22:15:02.340057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:15:02.340057Z digest=sha256:9966026075b5f1499d6ea5e6f69471aca0f8d26fee7881b9b2b2cf2705c8cb81

Observation 260f516a-eb60-4a70-aa94-5a542647b211 · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T22:15:02.450017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:15:02.450017Z digest=sha256:2c931f2af6bf08c0ed1afac3408345415549143e8aaf3e657a57aa5c90532ba5

Observation d7ed91ed-796f-4940-b039-802d573b9f48 · outbound

This paper cites Categorical Reparameterization with Gumbel-Softmax.

Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs Categorical Reparameterization with Gumbel-Softmax

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T22:15:02.527447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:15:02.527447Z digest=sha256:0858aae0940d711b318deec6e9a024779804096c152cb386e291976879869b85

Observation e1403065-aac9-4fcd-bfdc-403074e446ef · outbound

This paper cites Chat-univi: Unified visual representation em- powers large language models with image and video un- derstanding.

Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs Chat-univi: Unified visual representation em- powers large language models with image and video un- derstanding

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:15:08.677569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:15:02.615899Z digest=sha256:eef35a0301f51e54d39cfc9789bfdb417f1df61068bdde3e2990855d1eaa7c27

Observation 49eacf54-21f3-4a01-b6c3-b164f5b11144 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs LLaVA-OneVision: Easy Visual Task Transfer

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T22:15:02.714570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:15:02.714570Z digest=sha256:6b870ed8ffb02d8e93dfd6a996a0deac5f2f873791b3675cb1e3f5dede957ad4

Observation f53e097e-80e5-4101-b645-9a7a49a929b7 · outbound

This paper cites Mvbench: A comprehensive multi-modal video understand- ing benchmark.

Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs Mvbench: A comprehensive multi-modal video understand- ing benchmark

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:15:08.537709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:15:02.778441Z digest=sha256:3b2754d3b065bdd04c16e290775dd47748dabc8fa5616881faa69c47e560b30b

Observation a797c863-8993-4d59-b303-c2213ace2331 · outbound

This paper cites Mon- key: Image resolution and text label are important things for large multi-modal models.

Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs Mon- key: Image resolution and text label are important things for large multi-modal models

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:15:08.389761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:15:02.859142Z digest=sha256:1fecd5df3cffe3e1f1da1185893e9c0b190a0a7966613efbfa32f23d09782c54

Observation 76dc95d8-f7e4-4b88-9a02-94e10a7c9b77 · outbound

This paper cites KeyVideoLLM: Towards Large-scale Video Keyframe Selection.

Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs KeyVideoLLM: Towards Large-scale Video Keyframe Selection

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T22:15:02.922173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:15:02.922173Z digest=sha256:ae3d0c6a2d69b0f3e9eb2750587cfd3ad1e72c53af1883b5408b6ecb59099581

Observation b2dd8b82-c552-47d0-885a-c4f771f400d5 · outbound

This paper cites Video-llava: Learning united visual representation by alignment before projection.

Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs Video-llava: Learning united visual representation by alignment before projection

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:15:08.243146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:15:02.991356Z digest=sha256:3cb39082a9caeac1a8086857b3c9ba76ba3217454d1ccb319820f651c4a28380

Observation 2475e21a-6cd2-4261-8e08-9faca5ee815c · outbound

This paper cites Vila: On pre-training for vi- sual language models.

Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs Vila: On pre-training for vi- sual language models

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:15:08.048504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:15:03.062699Z digest=sha256:ae57e0f9c642fc429c4c0ef2cd21b202b6b95fbcadaed43762ab5eb174a4db19

Observation d5d2932f-6139-47ee-8d33-e1b0201e24e6 · outbound

This paper cites Visual instruction tuning.

Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs Visual instruction tuning

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:15:07.838909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:15:03.144052Z digest=sha256:0d1f8614472cf3101aa90229b6d660655fb5e705477cbe92bd66cb16bf382093

Observation 73badf97-5289-4ab2-9a12-ae2e30d2d29d · outbound

This paper cites Improved baselines with visual instruction tuning.

Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs Improved baselines with visual instruction tuning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T22:15:03.202524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:15:03.202524Z digest=sha256:069d3751c41c43add2dde297d96351f538c5f051dae369a1272c018ffa409968

Observation b8379fdf-f320-447c-9623-29220e3249d2 · outbound

This paper cites Llavanext: Improved reasoning, ocr, and world knowledge, 2024.

Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs Llavanext: Improved reasoning, ocr, and world knowledge, 2024

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:15:07.653968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:15:03.286726Z digest=sha256:9044ac0b1e50a7f0d9e4bbd38225dd6df97cfd81a2f7ad506348667525cc4db8

Observation 0f33d806-78f9-4ccf-9bd3-34caa259491c · outbound

This paper cites Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input.

Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T22:15:03.351183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:15:03.351183Z digest=sha256:4a437cf645a20083cb68b3c4243890da0f4409b546a1972aad01d0b23f9a5b2a

Observation 666db4a5-e386-4eac-b8a7-5681a16d0392 · outbound

This paper cites Video-chatgpt: Towards detailed video understanding via large vision and language models.

Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs Video-chatgpt: Towards detailed video understanding via large vision and language models

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:15:07.490542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:15:03.417538Z digest=sha256:4011f7f0ace14b7a4f1091f75cf75add93c03578faba4f4863f6ea28814f8663

Observation ad62d436-89cf-4000-b98e-3e65e6ed29ca · outbound

This paper cites Hello gpt-4o, 2024.

Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs Hello gpt-4o, 2024

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:15:07.287773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:15:03.481220Z digest=sha256:2ee1129a90bbe1e9bbc7bd672395e32187b7213d4baa650466da9fc35fa3a3f8

Observation 9fbe48c4-ba7c-415b-a164-a6f7af190310 · outbound

This paper cites Learn- ing transferable visual models from natural language super- vision.

Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs Learn- ing transferable visual models from natural language super- vision

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:15:07.110497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:15:03.561078Z digest=sha256:894aa776cf4b626091ad416c1707a7cbd9fcf2d84267c775f9e3c91117d3f4db

Observation 20018ddc-e8da-4d12-8dd6-b41804ca24cd · outbound

This paper cites Timechat: A time-sensitive multimodal large lan- guage model for long video understanding.

Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs Timechat: A time-sensitive multimodal large lan- guage model for long video understanding

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:15:06.906491Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:15:03.678869Z digest=sha256:32774cd0d5bee121b6d83e467d390ab2c3c29ef2caa762d1af412315f9d9bb7d

Observation 51b1b8f9-be18-4d35-bc41-a649916768d0 · outbound

This paper cites Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding.

Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T22:15:03.755141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:15:03.755141Z digest=sha256:168038c62d30b8936bfd1cf0eacf87d7546cc8d3e649bd795a65cf69591bfb89

Observation 5f3b6777-e92f-4325-91a5-bf8b25fc6660 · outbound

This paper cites Moviechat: From dense token to sparse memory for long video understanding.

Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs Moviechat: From dense token to sparse memory for long video understanding

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T22:15:03.824422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:15:03.824422Z digest=sha256:f7c8333e5997e4b787590a7fa22f076dd48b30c82997b2097384cc5b7235e51a

Observation b47ab2da-1955-449d-998b-88df156c4469 · outbound

This paper cites Video understanding with large language models: A survey.

Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs Video understanding with large language models: A survey

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T22:15:03.875000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:15:03.875000Z digest=sha256:76b15141a170e78d04940e46dbb67fd3a5760e3e94429b315c069a156c63f381

Observation e4359f59-bc8b-46fe-a718-628a619f961b · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T22:15:03.971684Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:15:03.971684Z digest=sha256:a7624a9f4be9d3cd9c3f4c2875458d2205659c61ba8376dec479990847bdb8c9

Observation 1f4c3484-0531-472f-ad6f-3aa4f348d22f · outbound

This paper cites Internvideo2: Scaling foundation models for multi- modal video understanding.

Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs Internvideo2: Scaling foundation models for multi- modal video understanding

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:15:06.755357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:15:04.029772Z digest=sha256:d59e02457233cfa5edf2d731bd8f3559efd2803c1795ffda1a8d7c5d68036859

Observation e39c4189-926f-46bb-b906-7f1542fec27a · outbound

This paper cites Longvideobench: A benchmark for long-context interleaved video-language understanding.

Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs Longvideobench: A benchmark for long-context interleaved video-language understanding

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:15:06.629724Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:15:04.102916Z digest=sha256:133b640b6c772fb32d60853250ec1e9287a2298c4d23308fc81d1b71b63ee9ed

Observation e73ebea1-c1cf-4f09-8858-be514963c3a3 · outbound

This paper cites Multi-agent reinforcement learning based frame sampling for effective untrimmed video recognition.

Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs Multi-agent reinforcement learning based frame sampling for effective untrimmed video recognition

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:15:06.502690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:15:04.241569Z digest=sha256:4e28622236dbefa7715fb3ca65fd0626ba141048ecd8d44eff4a6e65d347e2c6

Observation d916726f-0327-4842-ae41-4fcc2571c33c · outbound

This paper cites Adaframe: Adaptive frame selection for fast video recognition.

Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs Adaframe: Adaptive frame selection for fast video recognition

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:15:06.387469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:15:04.299593Z digest=sha256:22311c069ff12b8e0ba2bac1712b37de800a6c274a352a3850452679c16c4909

Observation a4b2f519-101d-44cc-af46-640da2b1c2df · outbound

This paper cites LongVILA: Scaling Long-Context Visual Language Models for Long Videos.

Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs LongVILA: Scaling Long-Context Visual Language Models for Long Videos

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T22:15:04.382683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:15:04.382683Z digest=sha256:0557c26a21079597a7a0c7655a13c543e3a867774e7ee5cee0dd2f8755e8f81f

Observation 665bd746-852e-4a8b-a786-607ecfbe37a9 · outbound

This paper cites Frame-voyager: Learning to query frames for video large language models.

Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs Frame-voyager: Learning to query frames for video large language models

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:15:06.261410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:15:04.480606Z digest=sha256:bcaffe0890c8f7a7c19e2b9331951c1c2aed1cc3fafb4729416f2af6aea0fe87

Observation eb4d00fa-9cc6-44eb-bf77-de94c0c3682b · outbound

This paper cites Sigmoid loss for language image pre-training.

Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs Sigmoid loss for language image pre-training

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:15:06.143689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:15:04.561512Z digest=sha256:ca1fd13f57369758e0bcf1bf027c2a3cebf1719f6f17f23686a8e5cb9dceb3e2

Observation e0e680ee-9a99-4961-857a-fd21ac5307dd · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T22:15:04.623326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:15:04.623326Z digest=sha256:ec5e05b611f73c4433959e9008680ab083e9abac8da31debfe3e2871a8710f27

Observation 5d985fbf-4b73-481f-b255-2b737e68829e · outbound

This paper cites LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models.

Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T22:15:04.686714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:15:04.686714Z digest=sha256:094d5f56c25bcd5d60eb399c7f305a633507b89b5738b1e9fd8a69d39993bae9

Observation ef8de3d9-0280-4974-bb2f-47853121def6 · outbound

This paper cites Long Context Transfer from Language to Vision.

Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs Long Context Transfer from Language to Vision

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T22:15:04.749344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:15:04.749344Z digest=sha256:407db9b63b22867b1b1a53c90e58258f6d5ae99d59eae26f2af70efa051dc8d3

Observation 5a1f0d10-f715-47c1-b335-f47ffddde9c3 · outbound

This paper cites Direct Preference Optimization of Video Large Multimodal Models from Language Model Reward.

Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs Direct Preference Optimization of Video Large Multimodal Models from Language Model Reward

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T22:15:04.864133Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:15:04.864133Z digest=sha256:f6235e259cb25dbae67a74cd6649edf88b8065976ecf949ff24614112918a9d3

Observation 0c307f2e-7cdd-4c08-b992-c4fc22cea03a · outbound

This paper cites LLaVA-Video: Video Instruction Tuning With Synthetic Data.

Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T22:15:05.019965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:15:05.019965Z digest=sha256:8c06ea99e2c85f4fb031aa0cd68130175430465da06d829b197423e7032d626a

Observation 9fc18216-6abf-4dc7-915d-f115ba96b59f · outbound

This paper cites Mgsampler: An explainable sampling strategy for video ac- tion recognition.

Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs Mgsampler: An explainable sampling strategy for video ac- tion recognition

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:15:05.983536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:15:05.086874Z digest=sha256:16190c9d697d616f46494b9c703261bec356921b8001d40cef68f21431f23d17

Observation f71645bf-2cfe-46ad-8924-4ae73dd3d962 · outbound

This paper cites MLVU: Benchmarking Multi-task Long Video Understanding.

Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs MLVU: Benchmarking Multi-task Long Video Understanding

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T22:15:05.172214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:15:05.172214Z digest=sha256:1dbe2bf5827481b85b9c24401b52563a12561a716086529a5ced1d4803b23605

Observation 2eef322c-e13d-4329-90c3-8916f65c57a9 · outbound

This paper cites Limitations Q-Frame enhances query-aware video understanding, but it depends on pre-trained models, lacks explicit temporal modeling, and operates within a fixed token budget.

Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs Limitations Q-Frame enhances query-aware video understanding, but it depends on pre-trained models, lacks explicit temporal modeling, and operates within a fixed token budget

Reference 2024

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:15:05.867099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:15:05.248695Z digest=sha256:ac51341cb888da86d13f2fc932a5bbe65b4a308190f0bce77845ab4027139935

Pith citing papers

Observation 0b6ad907-50e4-4dc3-bfa1-ecd47740152c · inbound

AdaSpark: Adaptive Sparsity for Efficient Long-Video Understanding cites this paper.

AdaSpark: Adaptive Sparsity for Efficient Long-Video Understanding Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:56:04.269758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T17:21:47.439019Z digest=sha256:ce25bc6e3c48359734af31a5c7ae37add474e9feeb290c1bedf8055410e1e50e

Observation 33844b79-8431-4eda-bdce-8ca1e28b38dc · inbound

Multi-Scale Separable Fourier Neural Networks for Solving High-Frequency PDEs cites this paper.

Multi-Scale Separable Fourier Neural Networks for Solving High-Frequency PDEs Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs

Reference 38

Resolution
unresolved
no resolver link, observed 2026-07-12T15:34:37.002001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T15:34:37.002001Z digest=sha256:a247887827269cb66f3c4abe11bc0ed7c72012cef2b4c81f6788e0dae8de9d6a

Observation eacdb4a2-0578-4ee1-87d6-3cc48c57148e · inbound

Multi-Scale Separable Fourier Neural Networks for Solving High-Frequency PDEs cites this paper.

Multi-Scale Separable Fourier Neural Networks for Solving High-Frequency PDEs Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs

Reference 38

Resolution
unresolved
no resolver link, observed 2026-07-14T18:39:24.915547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T18:39:24.915547Z digest=sha256:b5d1489c8da725b0a542f801549d39db1694f70ef97b3c8733c58a77eaed0ee2

Observation 079f6c34-65d1-42bd-aab7-82261c6a427e · inbound

PEEK: Picking Essential frames via Efficient Knowledge distillation cites this paper.

PEEK: Picking Essential frames via Efficient Knowledge distillation Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-07-01T19:16:01.318490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T22:46:53.704892Z digest=sha256:e9ff64ce0148cb87bcc3c9b8e5283b07bc8339cc5ed180b7e34c846cc0710b27

Observation 9498d452-63c8-44f2-8693-0f8a72695bb3 · inbound

Zero-Shot 3D Question Answering via Hierarchical View-to-Token Transportation cites this paper.

Zero-Shot 3D Question Answering via Hierarchical View-to-Token Transportation Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs

Reference 26

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T02:26:27.027565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-28T10:54:02.188634Z digest=sha256:2c5cd5ded7be3376c869636d58606410d927afc8cf5da8102e6a9bf8aa746db6

Observation de4ae604-e237-4b7a-9565-2c18234cc729 · inbound

Watch, Remember, Reason: Human-View Video Understanding with MLLMs cites this paper.

Watch, Remember, Reason: Human-View Video Understanding with MLLMs Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:27:15.057718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T22:00:28.350003Z digest=sha256:6b88e71c9c414a2b51e40342ce79595684e4f6f7047afa9983c478cfcd26a68f

Observation 9824f3dc-f80b-4c09-b06c-904f35ade8c6 · inbound

Towards Fast and Effective Long Video Understanding of Multimodal Large Language Models via Adaptive Quasi-Gaussian Sampling cites this paper.

Towards Fast and Effective Long Video Understanding of Multimodal Large Language Models via Adaptive Quasi-Gaussian Sampling Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-07-04T16:09:57.150830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-26T00:53:43.629684Z digest=sha256:ccde24d6a7736a0a5d78ce93ee1c56dacc8da6d0e618cc7a8f675135b8eb9b71

Observation 5051ba87-dc57-4d7a-946d-cbf469c36426 · inbound

QCA: Query- and Content-Aware Keyframe Selection for Long Video Understanding cites this paper.

QCA: Query- and Content-Aware Keyframe Selection for Long Video Understanding Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs

Reference 38

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T14:17:02.668832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-02T14:09:54.549499Z digest=sha256:1d7295d415affae9bc8cf9451a35db7520cd776f75630c728ef302c19cbacb9e

Observation 83ce1718-bef9-4d79-a24a-13456167b10f · inbound

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding cites this paper.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.137666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.137666Z digest=sha256:b12413bfa5859ad0e53c947c16a9645af4d5c2adc9b2d20171e5bbaeb18c3f2d