Pith. sign in

Paper Citation Record · LEDGER

CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding

As of 12 August 2026, this Paper Citation Record lists 21 of 21 outbound references and 26 inbound Pith citation observations for arXiv:2412.12075.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.12075 v1

Coverage vector

measured 21 of 21 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T14:22:51.060258Z

measured 47 of 47 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 26 of 26 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T00:43:37.679253Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T00:49:17.542022Z

Reference resolution

21 of 21 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved21
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5572d0dc-b182-4137-bdce-d2805dd1e64d · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding Gemini: A Family of Highly Capable Multimodal Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T14:22:50.946003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:22:50.946003Z digest=sha256:dc7827154cdb77c7b91f9821057123e2ffeea8f188523a9d9d3af9022c67fa07

Observation d3cf6368-4e3f-428a-bd60-eaf980c70b67 · outbound

This paper cites ReXTime: A Benchmark Suite for Reasoning-Across-Time in Videos.

CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding ReXTime: A Benchmark Suite for Reasoning-Across-Time in Videos

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T14:22:50.957256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:22:50.957256Z digest=sha256:6d28d9f6488eea03ea27be5d2b0dcc9209188f3489e6860450fe8b2a71cdd8d8

Observation daa6b931-4dae-412c-ad40-03b093044f16 · outbound

This paper cites How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites.

CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T14:22:50.963034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:22:50.963034Z digest=sha256:b4580b7e37fd72efee3181117253aa56aef6bde7abbb114ae2dcaf1c401389b5

Observation f63edbc4-30f1-4295-a97a-ba733a4ba18f · outbound

This paper cites Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos.

CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T14:22:50.974357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:22:50.974357Z digest=sha256:d452fa9431f7f215aab3c4b719df3ed1bfe988a6c870db50d262aaca5571a573

Observation 09751473-8f29-499f-bcff-c6650dc119af · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T14:22:50.980100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:22:50.980100Z digest=sha256:478b01b97ce113dc4762e3dd57a0be660e1213d2c40c13e2b851f08e33ba25cc

Observation 3961459f-b42d-4c2b-904b-f6c61d894f30 · outbound

This paper cites TVQA: Localized, Compositional Video Question Answering.

CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding TVQA: Localized, Compositional Video Question Answering

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T14:22:50.985707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:22:50.985707Z digest=sha256:e0144ed50dfdae45b07522015440e421a9a458b22595e66440787b9130e822be

Observation f5d29f87-55f5-4b11-b0c7-73eeaccb2ed3 · outbound

This paper cites VideoChat: Chat-Centric Video Understanding.

CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding VideoChat: Chat-Centric Video Understanding

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T14:22:50.996333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:22:50.996333Z digest=sha256:0c78e13352e198e714239c064f9b687cf13bb5da593937e3606d119ccb240467

Observation ceaf6c78-9c05-4091-8cc0-a2bd2c06ba25 · outbound

This paper cites Improved Baselines with Visual Instruction Tuning.

CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding Improved Baselines with Visual Instruction Tuning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T14:22:51.001598Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:22:51.001598Z digest=sha256:1351c60f95f33ce43057a1469c9710a54073a8bddda706f411abef1aae4ab047

Observation 1ad0e2d1-bf2a-4d42-8477-199a3c4c9914 · outbound

This paper cites Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input.

CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T14:22:51.006992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:22:51.006992Z digest=sha256:af952b7c200fbe498f24a5cd1087278dac1d6102a68cbb44e3bafac2b35995f7

Observation 0d49752f-6432-41a9-b1a4-dcbdc4b942ad · outbound

This paper cites Video-Bench: A Comprehensive Benchmark and Toolkit for Evaluating Video-based Large Language Models.

CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding Video-Bench: A Comprehensive Benchmark and Toolkit for Evaluating Video-based Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T14:22:51.012539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:22:51.012539Z digest=sha256:dce844acbc831782f677e4531d65fd1bd78355104c2c906cf27051e56f1d37ca

Observation 3de094b4-c346-49fb-978a-a29965673799 · outbound

This paper cites EgoVideo: Exploring Egocentric Foundation Model and Downstream Adaptation.

CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding EgoVideo: Exploring Egocentric Foundation Model and Downstream Adaptation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T14:22:51.018555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:22:51.018555Z digest=sha256:a6e4c2c776f74ddcaaaa2356299e509f746f8ca0d36aba4fb879699c5ddac894

Observation b2af867e-7a80-4ec6-a3c4-aec5c6374b46 · outbound

This paper cites CinePile: A Long Video Question Answering Dataset and Benchmark.

CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding CinePile: A Long Video Question Answering Dataset and Benchmark

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T14:22:51.024494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:22:51.024494Z digest=sha256:be52806adc481cf60b15d039a42fb7f96d18ecff909bfb7337484049ea0b58c3

Observation c5820709-7195-4e36-a83e-f2ecfae8add5 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T14:22:51.029774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:22:51.029774Z digest=sha256:d602cce2e47fd869101aac070d2c67011df0731c6fc09415e136b64d16e06d4d

Observation b9587f3f-bdc5-4272-8752-2b9c817eaaea · outbound

This paper cites LongVILA: Scaling Long-Context Visual Language Models for Long Videos.

CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding LongVILA: Scaling Long-Context Visual Language Models for Long Videos

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T14:22:51.041669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:22:51.041669Z digest=sha256:0754427d7e1d866af8a1b9d199b399a436ed35f33a4109df962b5a2ca8103cd6

Observation 1116ab4d-0408-4fce-8208-1804b28f0e7d · outbound

This paper cites MiniCPM-V: A GPT-4V Level MLLM on Your Phone.

CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T14:22:51.048774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:22:51.048774Z digest=sha256:aacc56fa7c0df3e89b50ae255dd0cad15dbd01251f205b134edcb1f5099fc6f0

Observation 091b6b9b-11a4-4835-a72c-8ac4d093758c · outbound

This paper cites MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities.

CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T14:22:51.055099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:22:51.055099Z digest=sha256:ba4e733ad6417551d9bda1ea4a6308b3141dda6844e54826255d9d68de6dc97c

Observation 8d64fd37-344d-4fb0-ba99-12c2337a4ef3 · outbound

This paper cites LvBench: A Benchmark for Long-form Video Understanding with Versatile Multi-modal Question Answering.

CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding LvBench: A Benchmark for Long-form Video Understanding with Versatile Multi-modal Question Answering

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T14:22:51.060258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:22:51.060258Z digest=sha256:d5886ffbc9fe02cfaabc26d9a1c8239cf5489b426c35c51981dcb0b305cf7794

Observation 312c4651-ca5f-49e3-aa59-2085c7554e85 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding LLaVA-OneVision: Easy Visual Task Transfer

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-11T14:22:50.990785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:22:50.990785Z digest=sha256:5f6c60a4da1f65650e93aaddd0b89f66f4560766f7f78db2472c4cc84317f216

Observation 0d86ed82-9362-493e-a91e-2c0729ec98b0 · outbound

This paper cites InternVideo2: Scaling Foundation Models for Multimodal Video Understanding.

CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding InternVideo2: Scaling Foundation Models for Multimodal Video Understanding

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-11T14:22:51.035972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:22:51.035972Z digest=sha256:64c596d5e0f27a7ec51c20c45249f517c35836fee9257892a89dbfd32797401a

Observation 13e58c84-453f-446d-aa6e-b6ad2e415696 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-11T14:22:50.951914Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:22:50.951914Z digest=sha256:3e2081210312973c3d22fef7e3036702880d3de2cd1079e596c2b8bf0f4d34f5

Observation c4427d20-f72d-465b-945a-e268ff6e010c · outbound

This paper cites MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding.

CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-11T14:22:50.968851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:22:50.968851Z digest=sha256:35d3711cb2ea91234b09f32c96a741d10261ead20643d57ddb999b327d570263

Pith citing papers

Observation c6578085-a73b-4727-b924-867d9c0aec1c · inbound

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model cites this paper.

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T23:07:14.381559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:07:14.381559Z digest=sha256:a3a86baa2db8a5b66b3c838600ceb0d9961dbed141ea26d4d7279350bec97b2a

Observation 6fedff7f-aaec-4ec1-8767-dea3b12c325d · inbound

MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment Grounding cites this paper.

MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment Grounding CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-19T13:17:18.536691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-19T13:13:40.485342Z digest=sha256:31c480e8827e6c16fe56c227f29dafc21cf398862cf6a722602baad1a142f550

Observation e150a606-02fb-4a4e-8b6f-c23c42111155 · inbound

DisTime: Distribution-based Time Representation for Video Large Language Models cites this paper.

DisTime: Distribution-based Time Representation for Video Large Language Models CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T12:32:50.149539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:32:50.149539Z digest=sha256:4a2dc8250b88b59d88c42212eb3600d1c9aadbdfcd7dd4fd13198d5168e1b62e

Observation 02223deb-a7f6-4379-8ff1-a37f4262b12d · inbound

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs cites this paper.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:05.529740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:05.529740Z digest=sha256:7efe449eaeb0fabf4105c4f8d2e2c839711b8665e9ed9a4dc898777f25bca13e

Observation e150c5ae-358f-4b63-a745-006588ae60ac · inbound

Bridging Perspectives: A Survey on Cross-view Collaborative Intelligence with Egocentric-Exocentric Vision cites this paper.

Bridging Perspectives: A Survey on Cross-view Collaborative Intelligence with Egocentric-Exocentric Vision CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:58.009101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:58.009101Z digest=sha256:eb0f01bd97aa134d111221f773108e745109639cc4f948d25a309e92e6e883b8

Observation d426b64f-a31c-4848-bfb2-48061f7ec977 · inbound

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding cites this paper.

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:56.855381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:56.855381Z digest=sha256:8524e75cfb1befd84f96ded1f702ea8f4241282b44fae2982b77dac6cc9484f1

Observation aa3c841a-826e-4063-b441-59d7aee0678b · inbound

VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos cites this paper.

VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T04:22:55.805315Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:22:55.805315Z digest=sha256:679a68801b9145ff20e9b39ff4f3db6e15c519bddf23c6779b0868f4937d7f44

Observation f0c1fecd-0d5e-46e2-96fd-1bf79aa69b5e · inbound

OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs cites this paper.

OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T22:41:07.070388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:41:07.070388Z digest=sha256:32d1c8be4fa10d8398d69fe5684019a7c0fb84b0b3d0020509d78b35d999a234

Observation 19e47eb9-e53e-4ef6-b09f-014ff4bff5ee · inbound

A Survey on Video Temporal Grounding with Multimodal Large Language Model cites this paper.

A Survey on Video Temporal Grounding with Multimodal Large Language Model CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-05T23:32:17.702112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:32:17.702112Z digest=sha256:93883afd5c7ccdd78b5b5eae137e8544fc055314ceffaf36568efdfe2de9f733

Observation 60dad2f5-2f7a-48c4-8107-13874d6f42b4 · inbound

EMCompress: Video-LLMs with Endomorphic Multimodal Compression cites this paper.

EMCompress: Video-LLMs with Endomorphic Multimodal Compression CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-18T20:41:50.596950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-18T20:40:42.995833Z digest=sha256:068c4a2d616aed189d8e664677de2570083d9ffea86e4c3c185ccb83210172e4

Observation 5b5217f4-7191-472c-8ce7-66cc9f303bd9 · inbound

REVISOR: Beyond Textual Reflection, Towards Multimodal Introspective Reasoning in Long-Form Video Understanding cites this paper.

REVISOR: Beyond Textual Reflection, Towards Multimodal Introspective Reasoning in Long-Form Video Understanding CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-17T22:20:22.941804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T22:19:36.366837Z digest=sha256:fdb52c1576ef44a97ad753971403a6e06f41a2b29a195f8e76e39ea470e7f90c

Observation df201b4e-fc59-42cd-b341-1f002533f6fa · inbound

VideoThinker: Building Agentic VideoLLMs with LLM-Guided Tool Reasoning cites this paper.

VideoThinker: Building Agentic VideoLLMs with LLM-Guided Tool Reasoning CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-16T12:17:51.923578Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T12:17:42.135851Z digest=sha256:c34deb9ef4978d2d2132a15648fd85c055312f59e1a4a485bf11e04defa480ac

Observation 8cb99da1-f860-46d7-964b-44a6e0bcaba9 · inbound

LongVideo-R1: Smart Navigation for Low-cost Long Video Understanding cites this paper.

LongVideo-R1: Smart Navigation for Low-cost Long Video Understanding CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:01:33.453527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-15T20:01:31.129959Z digest=sha256:c04da73d9b4e1f1616980189e7a3380bcf9fa927b2b6f83b2af43648790a9556

Observation 1e641f19-8de6-4e96-990a-0d80cdbc04a6 · inbound

Efficient Inference for Large Vision-Language Models: Bottlenecks, Techniques, and Prospects cites this paper.

Efficient Inference for Large Vision-Language Models: Bottlenecks, Techniques, and Prospects CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T23:45:50.955671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T18:54:04.104227Z digest=sha256:5439729b06f8b38268cbc691c61cc0de06c36a165759ff7b950e861e5e3c09e4

Observation 267929d3-e852-4c10-a1e2-0fc519398c0b · inbound

POINTS-Long: Adaptive Dual-Mode Visual Reasoning in MLLMs cites this paper.

POINTS-Long: Adaptive Dual-Mode Visual Reasoning in MLLMs CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:41:03.728687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T15:23:08.671342Z digest=sha256:543f73eb3d6a87cadbb9a2678c435918067eb315d15f9ffe92cf61603558f99a

Observation e610b167-f3da-415c-8fe0-d640e963b60a · inbound

Sink-Token-Aware Pruning for Fine-Grained Video Understanding in Efficient Video LLMs cites this paper.

Sink-Token-Aware Pruning for Fine-Grained Video Understanding in Efficient Video LLMs CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T00:44:48.414508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T00:43:44.921189Z digest=sha256:756f125f45bbcab3eedaf2d8d83530b7fadd3851432a19bf60ef3456500d3916

Observation 420a45c4-ccb8-4b28-ac77-fb6637033bd0 · inbound

Towards Temporal Compositional Reasoning in Long-Form Sports Videos cites this paper.

Towards Temporal Compositional Reasoning in Long-Form Sports Videos CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding

Reference 3

Resolution
unresolved
no resolver link, observed 2026-07-14T19:27:29.843866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T19:27:29.843866Z digest=sha256:0e398f9e78cd1e30cda0ab100ca77f36f72071c6781161ed41a374e5abce94b7

Observation bb6e4c7b-c7d6-46af-9ab0-a9de1463aa6a · inbound

EgoMemReason: A Memory-Driven Reasoning Benchmark for Long-Horizon Egocentric Video Understanding cites this paper.

EgoMemReason: A Memory-Driven Reasoning Benchmark for Long-Horizon Egocentric Video Understanding CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding

Reference 85

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T06:06:24.690741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-12T04:37:46.090777Z digest=sha256:a2353e77b04c922d26eaea3843bff474d59a27020def2ed79919cd50d7690622

Observation c8e94c6d-64a2-4df4-bab7-3e695f1db87e · inbound

VideoOdyssey: A Benchmark for Ultra-Long-Context and Omni-Modal Video Understanding cites this paper.

VideoOdyssey: A Benchmark for Ultra-Long-Context and Omni-Modal Video Understanding CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-25T05:55:25.014007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-25T05:51:49.390597Z digest=sha256:aa3aff8dc0ed728da601b8065b5efe00239156edb55a9d64acb88169a00a1a28

Observation 6aa79895-1600-4daf-9aec-21d29475123a · inbound

CaST-Bench: Benchmarking Causal Chain-Grounded Spatio-Temporal Reasoning for Video Question Answering cites this paper.

CaST-Bench: Benchmarking Causal Chain-Grounded Spatio-Temporal Reasoning for Video Question Answering CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-25T04:55:23.166760Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-25T04:54:23.077914Z digest=sha256:4235ce0e28ecbd7076d82d35173d0a345b2290ac9d990472e82aeb3f8bd04a1d

Observation 9c8f22a7-42e6-4db0-b5e9-7ccde862993c · inbound

CaST-Bench: Benchmarking Causal Chain-Grounded Spatio-Temporal Reasoning for Video Question Answering cites this paper.

CaST-Bench: Benchmarking Causal Chain-Grounded Spatio-Temporal Reasoning for Video Question Answering CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-07-04T00:49:17.545143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-07-04T00:41:02.284215Z digest=sha256:eee497388ee3f789321f97e299af6be023b3c3600fa80e4019d90479181dcd23

Observation f043ac01-df25-4a2e-86ed-07be12eda13a · inbound

Test-Time Scaling in Multimodal Foundation Models: A Comprehensive Survey of Generation and Reasoning cites this paper.

Test-Time Scaling in Multimodal Foundation Models: A Comprehensive Survey of Generation and Reasoning CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-07-02T21:37:25.333752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-06-27T19:36:57.231932Z digest=sha256:c1efcffca14c90b847a80648f82df7257fdf2d122f3b72093eca69fc39ace091

Observation 6d90e396-1fe9-4aa3-8d49-02c641377eef · inbound

Rethinking RAG in Long Videos: What to Retrieve and How to Use It? cites this paper.

Rethinking RAG in Long Videos: What to Retrieve and How to Use It? CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-07-03T15:28:33.999172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-27T06:30:33.428489Z digest=sha256:158e5aa7a8148f46a192a13aa3cdf15d6aac250bff0a4e0947b806496199eff7

Observation 432ebe94-bbaf-4073-84f0-3a6c78567a71 · inbound

Incentivizing Vision Language Models to Search for Long Video Question Answering cites this paper.

Incentivizing Vision Language Models to Search for Long Video Question Answering CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-12T05:50:16.895740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T05:50:16.895740Z digest=sha256:cbc25b9187e30fb205cb23474aa3083216971264b71dd4266071f3088b25ef93

Observation a86fe1e0-42ac-4ff6-a65f-c6d62ca74575 · inbound

TimeThink: Reasoning with Time for Video LLMs cites this paper.

TimeThink: Reasoning with Time for Video LLMs CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-11T08:59:46.244502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T08:59:46.244502Z digest=sha256:8cf8aed3eb3567c4f1b8b200d52dfabfd23df926f12a46754de6befbd6468b84

Observation 89ae2c35-e156-4fa2-8fdf-0a06d1aefbe6 · inbound

SCOUT: Self-Checking and Recovery-Aware Tool-Thought Agents for Ultra-Long Egocentric Video Reasoning cites this paper.

SCOUT: Self-Checking and Recovery-Aware Tool-Thought Agents for Ultra-Long Egocentric Video Reasoning CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T00:43:37.679253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:43:37.679253Z digest=sha256:76ce2209e61665d8f8bd1db786eea236ce4b425251b6cf25c80026e072135383