Pith. sign in

Paper Citation Record · LEDGER

FRAG: Frame Selection Augmented Generation for Long Video and Long Document Understanding

As of 21 August 2026, this Paper Citation Record lists 46 of 46 outbound references and 7 inbound Pith citation observations for arXiv:2504.17447.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.17447 v1

Coverage vector

measured 46 of 46 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T10:43:17.664668Z

measured 53 of 53 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T14:57:42.167776Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T16:09:57.135194Z

Reference resolution

46 of 46 outbound references displayed

  • verified exact0
  • verified fuzzy14
  • unresolved32
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1ada4143-35d9-45db-94cb-e14e07ef3683 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

FRAG: Frame Selection Augmented Generation for Long Video and Long Document Understanding Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-16T10:43:17.451471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:43:17.451471Z digest=sha256:5a88c802025e117d9a0cc6e1d23f01ccafdfd80e6e0c125fe734a9f221f428a6

Observation f89af446-a8ea-472b-903a-6491e0a3f9c1 · outbound

This paper cites Memory Consolidation Enables Long-Context Video Understanding.

FRAG: Frame Selection Augmented Generation for Long Video and Long Document Understanding Memory Consolidation Enables Long-Context Video Understanding

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-16T10:43:17.457003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:43:17.457003Z digest=sha256:a62f89c269e63fa707736dda9104d74f003b15826d3050eebad2be2b463e49e3

Observation 3b633a9e-24a4-4eea-b8d9-6ed4b62d17bd · outbound

This paper cites Scene text visual question answering.

FRAG: Frame Selection Augmented Generation for Long Video and Long Document Understanding Scene text visual question answering

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:43:18.359169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T10:43:17.462083Z digest=sha256:f45b472dc643820ef8b09dfee01ec92badd7d934798c969d59664f8ad7aecb0c

Observation e6059f84-e991-4470-9dc6-8990444182dd · outbound

This paper cites Gram: Global reasoning for multi-page vqa.

FRAG: Frame Selection Augmented Generation for Long Video and Long Document Understanding Gram: Global reasoning for multi-page vqa

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:43:18.344862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T10:43:17.466985Z digest=sha256:9ad34d16bb57044732a2236b11d72e13d25d4a95b539db25e63b247df572bb05

Observation 178e8cc1-86ac-4fb2-8dca-3d66c9b14994 · outbound

This paper cites How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites.

FRAG: Frame Selection Augmented Generation for Long Video and Long Document Understanding How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-16T10:43:17.471691Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:43:17.471691Z digest=sha256:fbba53c39a85564a790bbacbacd69b0065a38cc85b190bd75c073f96c1c0182e

Observation ce5fd533-6a47-45dd-95a7-9eeefb8dfe10 · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

FRAG: Frame Selection Augmented Generation for Long Video and Long Document Understanding VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-16T10:43:17.476787Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:43:17.476787Z digest=sha256:b2a48f641a82a0fc768d8d2b894d9fb0983be2c2e7818007cfad54bc7cb1bb5a

Observation c822df91-9e47-4ab0-ab21-1567456bcd32 · outbound

This paper cites InternLM-XComposer2-4KHD: A Pioneering Large Vision-Language Model Handling Resolutions from 336 Pixels to 4K HD.

FRAG: Frame Selection Augmented Generation for Long Video and Long Document Understanding InternLM-XComposer2-4KHD: A Pioneering Large Vision-Language Model Handling Resolutions from 336 Pixels to 4K HD

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-16T10:43:17.482676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:43:17.482676Z digest=sha256:10b179eadafc6b97378601ffcd61f95f02589f5a680bf71914648cf00b959a1d

Observation 6de5e810-6df3-4d82-b3ea-bae87cf764c5 · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

FRAG: Frame Selection Augmented Generation for Long Video and Long Document Understanding Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-16T10:43:17.487971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:43:17.487971Z digest=sha256:ae9ba92b9b88f6d3f129758e968ce7804945b5fdb0f2b659d59fbe758acaf1c7

Observation af9e289b-df83-4efe-a58f-38d089121ef5 · outbound

This paper cites Ma-lmm: Memory-augmented large multimodal model for long-term video understanding.

FRAG: Frame Selection Augmented Generation for Long Video and Long Document Understanding Ma-lmm: Memory-augmented large multimodal model for long-term video understanding

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:43:18.331002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T10:43:17.493137Z digest=sha256:74c9c7b89afdcd0357bf65887d9a7cc034f18e35ec78f53b7450330595fc7f1b

Observation 20baadda-17aa-49cc-bb4e-72a54323a851 · outbound

This paper cites mPLUG-DocOwl 1.5: Unified Structure Learning for OCR-free Document Understanding.

FRAG: Frame Selection Augmented Generation for Long Video and Long Document Understanding mPLUG-DocOwl 1.5: Unified Structure Learning for OCR-free Document Understanding

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-16T10:43:17.497736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:43:17.497736Z digest=sha256:09edf86606610061e96a0ea1c92d9aaada3f2c7205be9f001aebf8f23028fcbb

Observation e97737cd-0f35-4ae0-80d8-acfea68998ad · outbound

This paper cites mPLUG-DocOwl2: High-resolution Compressing for OCR-free Multi-page Document Understanding.

FRAG: Frame Selection Augmented Generation for Long Video and Long Document Understanding mPLUG-DocOwl2: High-resolution Compressing for OCR-free Multi-page Document Understanding

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-16T10:43:17.502876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:43:17.502876Z digest=sha256:924bb3f3feda10af821f46f00fb5cf5779387387acf0a98a766d95ac97c2b3e2

Observation 3c9a97cb-fba8-4e27-a9d7-e688c47c9bb3 · outbound

This paper cites Leopard: A Vision Language Model For Text-Rich Multi-Image Tasks.

FRAG: Frame Selection Augmented Generation for Long Video and Long Document Understanding Leopard: A Vision Language Model For Text-Rich Multi-Image Tasks

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-16T10:43:17.507757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:43:17.507757Z digest=sha256:395bc6746719b85ec373b87c47b5759c9acd502eea95929d85cb4532d012e245

Observation 5f6767b8-faa8-49bf-b35d-290f5056b60a · outbound

This paper cites Scsampler: Sampling salient clips from video for efficient action recog- nition.

FRAG: Frame Selection Augmented Generation for Long Video and Long Document Understanding Scsampler: Sampling salient clips from video for efficient action recog- nition

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:43:18.316744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T10:43:17.513237Z digest=sha256:b886eec3525d573fbf94b7c5ebc273e43e1af4440e103498f2e7f73b3f1ac2d6

Observation b86dcd42-eb52-46b0-8cce-528cb082ef19 · outbound

This paper cites Text-Conditioned Resampler For Long Form Video Understanding.

FRAG: Frame Selection Augmented Generation for Long Video and Long Document Understanding Text-Conditioned Resampler For Long Form Video Understanding

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-16T10:43:17.518068Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:43:17.518068Z digest=sha256:9865da99b96f6a38d19b8c66615e1af751dc889719eab48a7b877c04a0e1ee5f

Observation 36dc74f4-5747-4cbc-a24f-9190fe7a50ef · outbound

This paper cites Building and better understanding vision- language models: insights and future directions., 2024.

FRAG: Frame Selection Augmented Generation for Long Video and Long Document Understanding Building and better understanding vision- language models: insights and future directions., 2024

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:43:18.303019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T10:43:17.522825Z digest=sha256:4a7e2b17346fddd858e52d60cf4645b44aa3757ddd0e1e4cedfd674d5737a22b

Observation 58863500-c589-487f-90ce-db3d2db49149 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

FRAG: Frame Selection Augmented Generation for Long Video and Long Document Understanding LLaVA-OneVision: Easy Visual Task Transfer

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-16T10:43:17.527084Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:43:17.527084Z digest=sha256:bd05d6130a7218aa259134550143aeca8e9a41ff1df9471102724bdc6f0fb524

Observation 011b615d-ccd1-4494-8f3d-1ca725d6d4cb · outbound

This paper cites LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models.

FRAG: Frame Selection Augmented Generation for Long Video and Long Document Understanding LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-16T10:43:17.532350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:43:17.532350Z digest=sha256:03dc25eafac7bd2457600bddaeb1cf5ea9864fdcca5ffcfebff9a9c250bccdf3

Observation c1217fc4-556b-436b-9b8a-84ed79f83219 · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

FRAG: Frame Selection Augmented Generation for Long Video and Long Document Understanding Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-16T10:43:17.536688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:43:17.536688Z digest=sha256:03c436e4c0b4f0cf6fb36e74e5397213d902c0fb7864c7305c8506d8aa0b3b45

Observation 43381703-c98e-4d46-b49d-f46925d5cb91 · outbound

This paper cites Vila: On pre-training for visual language models.

FRAG: Frame Selection Augmented Generation for Long Video and Long Document Understanding Vila: On pre-training for visual language models

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:43:18.279282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T10:43:17.541436Z digest=sha256:82bcbe750eadfb460c7ec7d59dff462ad924baaad5504473495700dafb78e2e4

Observation bb7b0b27-0a16-4ee6-9ad6-9cfce0a83da2 · outbound

This paper cites Visual instruction tuning.

FRAG: Frame Selection Augmented Generation for Long Video and Long Document Understanding Visual instruction tuning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-16T10:43:17.546188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:43:17.546188Z digest=sha256:3b26b9e58efb94ca03f9bfc1f1aba983f85ad79235fc14c44f938e5bbd8567d4

Observation 4759bcae-9a95-4113-ac08-8a535f435ccb · outbound

This paper cites MMLongBench-Doc: Benchmarking Long-context Document Understanding with Visualizations.

FRAG: Frame Selection Augmented Generation for Long Video and Long Document Understanding MMLongBench-Doc: Benchmarking Long-context Document Understanding with Visualizations

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-16T10:43:17.550648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:43:17.550648Z digest=sha256:ffaaa55f9405fb019e71c8566c0e8620986b4b8bd4bab67647c5c77be1e99d67

Observation 402462e8-cc6e-4395-b2f9-27734b8d7c44 · outbound

This paper cites Egoschema: A diagnostic benchmark for very long- form video language understanding.

FRAG: Frame Selection Augmented Generation for Long Video and Long Document Understanding Egoschema: A diagnostic benchmark for very long- form video language understanding

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:43:18.254615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T10:43:17.555175Z digest=sha256:b5ac974fb4b939a0d69f16972cd3cc4526f104cf2dc66c80c3afc52d71be950b

Observation 4c7c1924-1e88-4e5e-b8fd-eb3bbd8c7894 · outbound

This paper cites Docvqa: A dataset for vqa on document images.

FRAG: Frame Selection Augmented Generation for Long Video and Long Document Understanding Docvqa: A dataset for vqa on document images

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-16T10:43:17.559536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:43:17.559536Z digest=sha256:29f714581c7149cec5b0eb3d1ca256f7dfc9e0d3f3a5f47441cc7884bdcb2945

Observation 39c4baba-695d-4bbe-b9c2-6f2680e3f6f3 · outbound

This paper cites Too many frames, not all useful: Efficient strategies for long- form video qa.

FRAG: Frame Selection Augmented Generation for Long Video and Long Document Understanding Too many frames, not all useful: Efficient strategies for long- form video qa

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-16T10:43:17.564016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:43:17.564016Z digest=sha256:0357fe718e0f8fd409ad067a96b4321c76c5d73a110fea14fb7efd6941d429b0

Observation 82fedbd9-d060-4a2a-be62-bceff9ec89ba · outbound

This paper cites Streaming Long Video Understanding with Large Language Models.

FRAG: Frame Selection Augmented Generation for Long Video and Long Document Understanding Streaming Long Video Understanding with Large Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-16T10:43:17.567967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:43:17.567967Z digest=sha256:557b91504b0688c3232cbb21d62facbcfc672053c136dc87b1d51db713960ae8

Observation f1eed15c-fea4-476d-bc42-759301067f22 · outbound

This paper cites LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding.

FRAG: Frame Selection Augmented Generation for Long Video and Long Document Understanding LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-16T10:43:17.572620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:43:17.572620Z digest=sha256:467ed68ab79713fc672e43e47162dfb7aa67298a0e49802ad9554319432e893b

Observation 1857db82-5077-4b20-9c24-09374eac1b50 · outbound

This paper cites Slidevqa: A dataset for document visual question answering on multiple images.

FRAG: Frame Selection Augmented Generation for Long Video and Long Document Understanding Slidevqa: A dataset for document visual question answering on multiple images

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:43:18.230330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T10:43:17.577304Z digest=sha256:188b5a92d3a9fd2533632ff0dc7cbfbfa39c60951510adfac1f9b06048404ef4

Observation d7b5f9fb-345a-4dae-acf8-1fe0e52fb72a · outbound

This paper cites Instructdoc: A dataset for zero-shot gener- alization of visual document understanding with instructions.

FRAG: Frame Selection Augmented Generation for Long Video and Long Document Understanding Instructdoc: A dataset for zero-shot gener- alization of visual document understanding with instructions

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:43:18.214048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T10:43:17.581262Z digest=sha256:cffef57a9f792575ff064a6ce898bd7b482326fe0f00b79c5f818fee79346e61

Observation dfc16086-dddb-4efb-ad8e-79989b8bfdc0 · outbound

This paper cites Hi- erarchical multimodal transformers for multipage docvqa.

FRAG: Frame Selection Augmented Generation for Long Video and Long Document Understanding Hi- erarchical multimodal transformers for multipage docvqa

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:43:18.199451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T10:43:17.585932Z digest=sha256:67c28a47fc9b0da85cc4c775ff64d340003f993a90acf92039306f7c6d6ab800

Observation 1298d1a8-b6fb-4a92-8886-b831c7296f81 · outbound

This paper cites Tarsier: Recipes for Training and Evaluating Large Video Description Models.

FRAG: Frame Selection Augmented Generation for Long Video and Long Document Understanding Tarsier: Recipes for Training and Evaluating Large Video Description Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-16T10:43:17.590383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:43:17.590383Z digest=sha256:46c61c41b4088966e25fda237bc3caef509177e66c63eee3afa4a7a6d28de926

Observation 58ae6ca0-c904-4f4e-ab8f-d3c7e6208b1d · outbound

This paper cites Lvbench: An extreme long video under- standing benchmark, 2024.

FRAG: Frame Selection Augmented Generation for Long Video and Long Document Understanding Lvbench: An extreme long video under- standing benchmark, 2024

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:43:18.184832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T10:43:17.595458Z digest=sha256:e19a64ed2f4dd90f7810973b197404200130ca46ec4b33024946f411e0f94695

Observation 5206eac9-1b0d-441b-943d-4ca90021443f · outbound

This paper cites VideoAgent: Long-form Video Understanding with Large Language Model as Agent.

FRAG: Frame Selection Augmented Generation for Long Video and Long Document Understanding VideoAgent: Long-form Video Understanding with Large Language Model as Agent

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-16T10:43:17.599746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:43:17.599746Z digest=sha256:db7f5f26db388d18e1638756e68da37c451f4b12f159a56aff4dce28e39bd928

Observation a83f2be9-a27a-4baf-acab-f113d7f31e9a · outbound

This paper cites VideoTree: Adaptive Tree-based Video Representation for LLM Reasoning on Long Videos.

FRAG: Frame Selection Augmented Generation for Long Video and Long Document Understanding VideoTree: Adaptive Tree-based Video Representation for LLM Reasoning on Long Videos

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-16T10:43:17.604526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:43:17.604526Z digest=sha256:c7adbe1ad2ca2975bf50c9f8060a7fdcd2ad1d22b1ad7ac94a8f8bf4ef7ed566

Observation 9b4584ac-7149-4d2c-8a41-910b8a38eea8 · outbound

This paper cites LongVLM: Efficient Long Video Understanding via Large Language Models.

FRAG: Frame Selection Augmented Generation for Long Video and Long Document Understanding LongVLM: Efficient Long Video Understanding via Large Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-16T10:43:17.609201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:43:17.609201Z digest=sha256:324acd0518643a63c93f9773fac0c85faaddb4c5618eff6e403c3fba63462a15

Observation 83dd7ddf-af62-4574-84b1-80f13ceefb27 · outbound

This paper cites LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding.

FRAG: Frame Selection Augmented Generation for Long Video and Long Document Understanding LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-16T10:43:17.613816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:43:17.613816Z digest=sha256:331f5421ad5e6473ccf0cbe2beb47d747c78edf288ecaaa554c3c93c1dd47dc2

Observation 354fad58-8c66-4189-8728-22959c02b8fa · outbound

This paper cites Visual Haystacks: A Vision-Centric Needle-In-A-Haystack Benchmark.

FRAG: Frame Selection Augmented Generation for Long Video and Long Document Understanding Visual Haystacks: A Vision-Centric Needle-In-A-Haystack Benchmark

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-16T10:43:17.618225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:43:17.618225Z digest=sha256:f6dcb2558131588526ba56a2e3faba31c46fc8dbe2692532a1608d5a1258fd25

Observation ef481088-fd65-43c7-9e6e-30081001d4aa · outbound

This paper cites Next-qa: Next phase of question-answering to explaining temporal actions.

FRAG: Frame Selection Augmented Generation for Long Video and Long Document Understanding Next-qa: Next phase of question-answering to explaining temporal actions

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:43:18.170973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T10:43:17.623636Z digest=sha256:4c0be0e55fa2097073ad5edf4565bfcde33bb035eb2161cdb4c5cc4f7b625d05

Observation 8c82477e-1992-47de-bbd0-7070ecd3c9bb · outbound

This paper cites PDF-WuKong: A Large Multimodal Model for Efficient Long PDF Reading with End-to-End Sparse Sampling.

FRAG: Frame Selection Augmented Generation for Long Video and Long Document Understanding PDF-WuKong: A Large Multimodal Model for Efficient Long PDF Reading with End-to-End Sparse Sampling

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-16T10:43:17.627836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:43:17.627836Z digest=sha256:6feda61aac8038514752e543601e7ff925b0108643b031651f04672ae51c8e0f

Observation 5d5ac8eb-d889-43ed-a7b2-f3af9282befb · outbound

This paper cites LongVILA: Scaling Long-Context Visual Language Models for Long Videos.

FRAG: Frame Selection Augmented Generation for Long Video and Long Document Understanding LongVILA: Scaling Long-Context Visual Language Models for Long Videos

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-16T10:43:17.633250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:43:17.633250Z digest=sha256:50bb8aff4a515bed28cac23cdf4434e7ebd2041f7b47960d2547aabb22eb4ad4

Observation c4f6d4a8-51e4-4acc-9cb6-8c8208e0830f · outbound

This paper cites Self-chained image-language model for video localization and question answering.

FRAG: Frame Selection Augmented Generation for Long Video and Long Document Understanding Self-chained image-language model for video localization and question answering

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:43:18.156588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T10:43:17.637791Z digest=sha256:6de6c0bb9d0be81e3306d669fe99205f2b0a14ace7f3947fad09e8237ec74b0f

Observation a1c67c44-1506-41e7-bf7e-73884a4d6389 · outbound

This paper cites Frame-Voyager: Learning to Query Frames for Video Large Language Models.

FRAG: Frame Selection Augmented Generation for Long Video and Long Document Understanding Frame-Voyager: Learning to Query Frames for Video Large Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-16T10:43:17.642133Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:43:17.642133Z digest=sha256:a5969cd40189b826719dd97f95ded4ce94b6c48d53b3da026992b98e2b477252

Observation 3d8254d3-2911-476a-afff-4abea8b596ce · outbound

This paper cites Sigmoid loss for language image pre-training.

FRAG: Frame Selection Augmented Generation for Long Video and Long Document Understanding Sigmoid loss for language image pre-training

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-16T10:43:17.646612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:43:17.646612Z digest=sha256:c7d1ef950f07ac906986b5d697b14f2b42e742cf20c85ce992d8124b3b5248c2

Observation f8a7fa0d-12d6-43e3-bddb-2f8f684dfb1c · outbound

This paper cites Long Context Transfer from Language to Vision.

FRAG: Frame Selection Augmented Generation for Long Video and Long Document Understanding Long Context Transfer from Language to Vision

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-16T10:43:17.650992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:43:17.650992Z digest=sha256:9bbdf876d4c41df4ab183956cf0660a83c81dc7459c2e1d66dbe0d6b08472f53

Observation 5d3a52f5-c8a9-468c-855a-24bd4205849f · outbound

This paper cites Llava- next: A strong zero-shot video understanding model, 2024.

FRAG: Frame Selection Augmented Generation for Long Video and Long Document Understanding Llava- next: A strong zero-shot video understanding model, 2024

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-16T10:43:17.655560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:43:17.655560Z digest=sha256:21e58d3d3383ff7739de8d133d7f22844a49a8c54ec97beabc36e588a7b311b9

Observation 272592ad-4daa-4e37-95db-74214230670d · outbound

This paper cites MLVU: Benchmarking Multi-task Long Video Understanding.

FRAG: Frame Selection Augmented Generation for Long Video and Long Document Understanding MLVU: Benchmarking Multi-task Long Video Understanding

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-16T10:43:17.660147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:43:17.660147Z digest=sha256:4d94676958bc2ff91dee1a23e8ff876a9db5a8ba01f8c89fa492a22092557647

Observation f38d4079-ca5a-428e-a6af-ee1209482b68 · outbound

This paper cites Selection Scoring Prompt Our prompt is based on the prompt used in SeViLA [40].

FRAG: Frame Selection Augmented Generation for Long Video and Long Document Understanding Selection Scoring Prompt Our prompt is based on the prompt used in SeViLA [40]

Reference 2024

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:43:18.122194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T10:43:17.664668Z digest=sha256:71f6b814cad8ed3efc33adf7c7c13148c26650834b8153d833fe3c9311875c5d

Pith citing papers

Observation 34d43881-d5b2-4eb7-a79c-a3532b767e37 · inbound

HumanPCR: Probing MLLM Capabilities in Diverse Human-Centric Scenes cites this paper.

HumanPCR: Probing MLLM Capabilities in Diverse Human-Centric Scenes FRAG: Frame Selection Augmented Generation for Long Video and Long Document Understanding

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-05T19:03:08.072560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:03:08.072560Z digest=sha256:41c729356a8ab1ea32ca3ddbc2ca1f0de73188844e20a0fdf93e57b7731c7950

Observation b25ee4a3-3132-4bde-b546-d701ac7253df · inbound

Training-free Uncertainty Guidance for Complex Visual Tasks with MLLMs cites this paper.

Training-free Uncertainty Guidance for Complex Visual Tasks with MLLMs FRAG: Frame Selection Augmented Generation for Long Video and Long Document Understanding

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-04T13:27:56.481199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:27:56.481199Z digest=sha256:77f9105a671e56fa95b784dfd3574fbfa78b3121d7521f31b6eb761b56f009e5

Observation 9928cae8-7e0d-41b4-9ec5-2f63929189fc · inbound

GridProbe: Posterior-Probing for Adaptive Test-Time Compute in Long-Video VLMs cites this paper.

GridProbe: Posterior-Probing for Adaptive Test-Time Compute in Long-Video VLMs FRAG: Frame Selection Augmented Generation for Long Video and Long Document Understanding

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:56:31.830293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-12T03:46:31.117768Z digest=sha256:b53e434cfc28f56479869f0a95b1ef41a9f2ebd4d6c19be6fc9404d4c6e48952

Observation d481f219-c623-41d4-8a26-96118247b5cb · inbound

Towards Fast and Effective Long Video Understanding of Multimodal Large Language Models via Adaptive Quasi-Gaussian Sampling cites this paper.

Towards Fast and Effective Long Video Understanding of Multimodal Large Language Models via Adaptive Quasi-Gaussian Sampling FRAG: Frame Selection Augmented Generation for Long Video and Long Document Understanding

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-07-04T16:09:57.137169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-26T00:53:43.629684Z digest=sha256:48354c88d0758f693ddb45a8b075237af402368f1776ef69d8f2b1736ea3f305

Observation 4065af9c-3783-4808-9ff1-13758e62168b · inbound

QCA: Query- and Content-Aware Keyframe Selection for Long Video Understanding cites this paper.

QCA: Query- and Content-Aware Keyframe Selection for Long Video Understanding FRAG: Frame Selection Augmented Generation for Long Video and Long Document Understanding

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T14:17:02.680402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-02T14:09:54.549499Z digest=sha256:22f58abae7482deb6945d222c86d0dd5b0fe43b4af1b1825db7fabdf9e49ccd9

Observation d598b017-5f0a-4bc8-9264-60f35e2e2fb8 · inbound

QSVideo: Query-Conditioned Semantic Temporal Retrieval for Video Understanding cites this paper.

QSVideo: Query-Conditioned Semantic Temporal Retrieval for Video Understanding FRAG: Frame Selection Augmented Generation for Long Video and Long Document Understanding

Reference 25

Resolution
unresolved
no resolver link, observed 2026-07-11T17:18:41.284513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T17:18:41.284513Z digest=sha256:53f87f3ecaad50b31a71f213413d0df30c46b230ef4b3de2a42943a080f23ac3

Observation 917e71ba-b271-41a7-b130-9ebdc8cb9808 · inbound

Caved or Convinced: Temporal Sampling Gates Claim Deference in Video Large Language Models cites this paper.

Caved or Convinced: Temporal Sampling Gates Claim Deference in Video Large Language Models FRAG: Frame Selection Augmented Generation for Long Video and Long Document Understanding

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:42.167776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:57:42.167776Z digest=sha256:98cad4527d7f04246916c4274496272bfd119f7210c4c04803c98bd7e2464add