Pith. sign in

Paper Citation Record · LEDGER

When and What: Diffusion-Grounded VideoLLM with Entity Aware Segmentation for Long Video Understanding

As of 17 August 2026, this Paper Citation Record lists 18 of 18 outbound references and 1 inbound Pith citation observation for arXiv:2508.15641.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.15641 v1

Coverage vector

measured 18 of 18 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T17:50:50.164729Z

measured 19 of 19 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-07T09:12:13.460169Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-12T09:46:27.899181Z

Reference resolution

18 of 18 outbound references displayed

  • verified exact0
  • verified fuzzy9
  • unresolved9
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 86d4c77b-4a91-45b4-be7f-d6c03b50dd93 · outbound

This paper cites Vtimellm: Empower llm to grasp video moments,.

When and What: Diffusion-Grounded VideoLLM with Entity Aware Segmentation for Long Video Understanding Vtimellm: Empower llm to grasp video moments,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:50:52.029281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T17:50:48.479947Z digest=sha256:602711df2b88a44539b4c3c315847a0f5e9752d7ef68743506e74024b7745710

Observation aefeeac5-90d6-4422-91d1-defef698cfa2 · outbound

This paper cites Videotree: Adaptive tree-based video representation for llm reasoning on long videos,.

When and What: Diffusion-Grounded VideoLLM with Entity Aware Segmentation for Long Video Understanding Videotree: Adaptive tree-based video representation for llm reasoning on long videos,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:50:51.781116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T17:50:48.511981Z digest=sha256:4739147af0da5e920b6d233996791f52d1e303660d40fc99233db8525ecf1507

Observation 364333ec-be1b-490d-a1e3-cdf078de5905 · outbound

This paper cites Vtg-llm: Integrating timestamp knowledge into video llms for enhanced video temporal grounding,.

When and What: Diffusion-Grounded VideoLLM with Entity Aware Segmentation for Long Video Understanding Vtg-llm: Integrating timestamp knowledge into video llms for enhanced video temporal grounding,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:50:51.583649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T17:50:48.594059Z digest=sha256:65160e1cd74c2e7f5c47fa4c0334f6d0cda02dbd79163de74fc87562c6743378

Observation f557661a-375b-4218-bfff-7d2641477f6e · outbound

This paper cites Flamingo: a visual language model for few-shot learning,.

When and What: Diffusion-Grounded VideoLLM with Entity Aware Segmentation for Long Video Understanding Flamingo: a visual language model for few-shot learning,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:50:51.377592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T17:50:48.678859Z digest=sha256:67b97455fd56c6ab5b6a9d25b4aad5ae60ea0d142c705c05e56e9fe3b6f91677

Observation 0039bdd4-8112-4d21-9d69-1396c1188891 · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

When and What: Diffusion-Grounded VideoLLM with Entity Aware Segmentation for Long Video Understanding Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-05T17:50:48.760200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:50:48.760200Z digest=sha256:a52fe16999b058963b12af8d6e7f1bf16bab09684753d945b36a84c86f99b8a1

Observation 2769dd94-fbfa-4156-8c93-84d1125e4928 · outbound

This paper cites Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models.

When and What: Diffusion-Grounded VideoLLM with Entity Aware Segmentation for Long Video Understanding Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T17:50:48.860892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:50:48.860892Z digest=sha256:9d5eadfb125a10a4da77550289567bb072591ea587194b0add38efa30b964768

Observation 4a68319b-ab34-403f-824d-f917f278a170 · outbound

This paper cites Timechat: A time-sensitive multimodal large language model for long video understanding,.

When and What: Diffusion-Grounded VideoLLM with Entity Aware Segmentation for Long Video Understanding Timechat: A time-sensitive multimodal large language model for long video understanding,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:50:51.190776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T17:50:48.962891Z digest=sha256:dfdf1d4611949594101a306ea045b52a5150bffa034ce206f04267d2f4d34472

Observation a9a7e9ff-10cd-4e27-b89f-d12e6896098a · outbound

This paper cites Grounded-VideoLLM: Sharpening Fine-grained Temporal Grounding in Video Large Language Models.

When and What: Diffusion-Grounded VideoLLM with Entity Aware Segmentation for Long Video Understanding Grounded-VideoLLM: Sharpening Fine-grained Temporal Grounding in Video Large Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T17:50:49.044330Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:50:49.044330Z digest=sha256:6107efaa6e7b4cd7562855698560740ab82ecd71f684343e5df593eba25e1d42

Observation 094dc931-5348-485d-9c67-e8218fcb5610 · outbound

This paper cites VideoGPT+: Integrating Image and Video Encoders for Enhanced Video Understanding.

When and What: Diffusion-Grounded VideoLLM with Entity Aware Segmentation for Long Video Understanding VideoGPT+: Integrating Image and Video Encoders for Enhanced Video Understanding

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T17:50:49.173346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:50:49.173346Z digest=sha256:8e849188d6baf369c6360efdbe68e3d38b5f4006a66ccdf94533dbe4724672b5

Observation be38c7f1-2134-4f61-a717-636175509c2f · outbound

This paper cites Longvlm: Efficient long video understanding via large language models,.

When and What: Diffusion-Grounded VideoLLM with Entity Aware Segmentation for Long Video Understanding Longvlm: Efficient long video understanding via large language models,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:50:51.011850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T17:50:49.247131Z digest=sha256:52cccd2522e915a1677a7f8567eb2982c52989a816e0c2c15251c728b39a4b4c

Observation 627e59f1-c3ca-40f8-aa2d-712cb0a35502 · outbound

This paper cites Video summarization using denoising diffusion probabilistic model,.

When and What: Diffusion-Grounded VideoLLM with Entity Aware Segmentation for Long Video Understanding Video summarization using denoising diffusion probabilistic model,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:50:50.847614Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T17:50:49.390336Z digest=sha256:de930e069be70fa8fe596bb2d0a23bc8b5778d0af29732ad59808dd4c993b74f

Observation 9577ab86-3d90-4017-a53a-bb418e43127c · outbound

This paper cites Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos.

When and What: Diffusion-Grounded VideoLLM with Entity Aware Segmentation for Long Video Understanding Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T17:50:49.465440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:50:49.465440Z digest=sha256:8cf39bda50b7b1750fbd731da60695799a08e7637159655709d993e70c471a89

Observation f9154289-0194-427f-acab-353318a02020 · outbound

This paper cites Videoglamm: A large multimodal model for pixel-level visual grounding in videos,.

When and What: Diffusion-Grounded VideoLLM with Entity Aware Segmentation for Long Video Understanding Videoglamm: A large multimodal model for pixel-level visual grounding in videos,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:50:50.709653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T17:50:49.612907Z digest=sha256:4e7b36f4711358dc99100b10d291c2a2b3761cc35c80355b749bd5c8e1b44749

Observation e4090a78-d4c2-491d-ae9d-d464582bf38d · outbound

This paper cites VACE: All-in-One Video Creation and Editing.

When and What: Diffusion-Grounded VideoLLM with Entity Aware Segmentation for Long Video Understanding VACE: All-in-One Video Creation and Editing

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T17:50:49.717507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:50:49.717507Z digest=sha256:6f0371ec811e601177c0b664d0eede25579e3e5dbecd62f55f6ba4cc57491111

Observation 8ffbb9f7-fe06-45a1-8331-7f38b3832e13 · outbound

This paper cites Wan: Open and Advanced Large-Scale Video Generative Models.

When and What: Diffusion-Grounded VideoLLM with Entity Aware Segmentation for Long Video Understanding Wan: Open and Advanced Large-Scale Video Generative Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T17:50:49.870293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:50:49.870293Z digest=sha256:ebf0d875c134c8f0ef06f18d07597f51dd9881c260dc04813c995a367f0935e7

Observation ffe4b62a-7b5d-47bb-9c87-db8bf28ba84e · outbound

This paper cites Can i trust your answer? visually grounded video question answering,.

When and What: Diffusion-Grounded VideoLLM with Entity Aware Segmentation for Long Video Understanding Can i trust your answer? visually grounded video question answering,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:50:50.460962Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T17:50:50.011389Z digest=sha256:492345b30c09be70f8257431755407c29ce4844f41c3ac021238e10a75aafd8a

Observation c5f7169c-70dc-4323-abb9-7951d3639879 · outbound

This paper cites Phi-4 Technical Report.

When and What: Diffusion-Grounded VideoLLM with Entity Aware Segmentation for Long Video Understanding Phi-4 Technical Report

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T17:50:50.085949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:50:50.085949Z digest=sha256:733745f54335d28ebc47f9ed091896bd71916aafc7330c2014b992ea0ee95069

Observation 26c83433-9103-4c3f-a97f-767d35310425 · outbound

This paper cites Grounded SAM: Assembling Open-World Models for Diverse Visual Tasks.

When and What: Diffusion-Grounded VideoLLM with Entity Aware Segmentation for Long Video Understanding Grounded SAM: Assembling Open-World Models for Diverse Visual Tasks

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T17:50:50.164729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:50:50.164729Z digest=sha256:bb20c5ecf60181dddc454279dfcb8c2a29f31b05c77cacf2c6b3512bc9fac6d2

Pith citing papers

Observation 5238bb53-07cb-4d26-88e8-8d378f7169a6 · inbound

YOSE: You Only Select Essential Tokens for Efficient DiT-based Video Object Removal cites this paper.

YOSE: You Only Select Essential Tokens for Efficient DiT-based Video Object Removal When and What: Diffusion-Grounded VideoLLM with Entity Aware Segmentation for Long Video Understanding

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:46:27.901967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-07T09:12:13.460169Z digest=sha256:c56fc4dd5865db06b399a7ebb7fe083bffd153835944598b03259aae5c93505f