Pith. sign in

Paper Citation Record · LEDGER

When and What: Diffusion-Grounded VideoLLM with Entity Aware Segmentation for Long Video Understanding

As of 9 August 2026, this Paper Citation Record lists 18 of 18 outbound references and 1 inbound Pith citation observation for arXiv:2508.15641.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.15641 v1

Coverage vector

measured 18 of 18 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T17:50:50.164729Z

measured 19 of 19 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-07T09:12:13.460169Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-12T09:46:27.899181Z

Reference resolution

18 of 18 outbound references displayed

  • verified exact0
  • verified fuzzy9
  • unresolved9
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 86d4c77b-4a91-45b4-be7f-d6c03b50dd93 · outbound

This paper cites Vtimellm: Empower llm to grasp video moments,.

When and What: Diffusion-Grounded VideoLLM with Entity Aware Segmentation for Long Video Understanding Vtimellm: Empower llm to grasp video moments,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:50:52.029281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T17:50:48.479947Z digest=sha256:b10fc017d56b336be463a4d3d825bc0ab09a6630354b546436cb9462343b7d67

Observation aefeeac5-90d6-4422-91d1-defef698cfa2 · outbound

This paper cites Videotree: Adaptive tree-based video representation for llm reasoning on long videos,.

When and What: Diffusion-Grounded VideoLLM with Entity Aware Segmentation for Long Video Understanding Videotree: Adaptive tree-based video representation for llm reasoning on long videos,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:50:51.781116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T17:50:48.511981Z digest=sha256:06b51e0c4a582888cf7923909110316306e088fc62d881fbc27b417d2d8a6b49

Observation 364333ec-be1b-490d-a1e3-cdf078de5905 · outbound

This paper cites Vtg-llm: Integrating timestamp knowledge into video llms for enhanced video temporal grounding,.

When and What: Diffusion-Grounded VideoLLM with Entity Aware Segmentation for Long Video Understanding Vtg-llm: Integrating timestamp knowledge into video llms for enhanced video temporal grounding,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:50:51.583649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T17:50:48.594059Z digest=sha256:1758b9919adc142ea4140d7ae2212813cc72a2c87ef84f55d831edeeceb05001

Observation f557661a-375b-4218-bfff-7d2641477f6e · outbound

This paper cites Flamingo: a visual language model for few-shot learning,.

When and What: Diffusion-Grounded VideoLLM with Entity Aware Segmentation for Long Video Understanding Flamingo: a visual language model for few-shot learning,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:50:51.377592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T17:50:48.678859Z digest=sha256:075f2112b79cb688447f7fbfe7437a04c1b40b4cc8689e8a9a750b107cb2e388

Observation 0039bdd4-8112-4d21-9d69-1396c1188891 · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

When and What: Diffusion-Grounded VideoLLM with Entity Aware Segmentation for Long Video Understanding Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-05T17:50:48.760200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:50:48.760200Z digest=sha256:ef395c252e6e17208d71077b467c53bffdd1253bd94750cad21f8271a6ac961a

Observation 2769dd94-fbfa-4156-8c93-84d1125e4928 · outbound

This paper cites Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models.

When and What: Diffusion-Grounded VideoLLM with Entity Aware Segmentation for Long Video Understanding Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T17:50:48.860892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:50:48.860892Z digest=sha256:ea1d9b7a50c733dfa900e29119edbf50e17a64ce9f0022d4e4c1dce606437ed7

Observation 4a68319b-ab34-403f-824d-f917f278a170 · outbound

This paper cites Timechat: A time-sensitive multimodal large language model for long video understanding,.

When and What: Diffusion-Grounded VideoLLM with Entity Aware Segmentation for Long Video Understanding Timechat: A time-sensitive multimodal large language model for long video understanding,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:50:51.190776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T17:50:48.962891Z digest=sha256:1cf4c8e66e434b8ac0b47b63f8173430f7990335d086c628b9a4d85f3616a090

Observation a9a7e9ff-10cd-4e27-b89f-d12e6896098a · outbound

This paper cites Grounded-VideoLLM: Sharpening Fine-grained Temporal Grounding in Video Large Language Models.

When and What: Diffusion-Grounded VideoLLM with Entity Aware Segmentation for Long Video Understanding Grounded-VideoLLM: Sharpening Fine-grained Temporal Grounding in Video Large Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T17:50:49.044330Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:50:49.044330Z digest=sha256:245e4bd66b7c40d242ab084271d3b7a3b3155053cf09d639261090682b2610f3

Observation 094dc931-5348-485d-9c67-e8218fcb5610 · outbound

This paper cites VideoGPT+: Integrating Image and Video Encoders for Enhanced Video Understanding.

When and What: Diffusion-Grounded VideoLLM with Entity Aware Segmentation for Long Video Understanding VideoGPT+: Integrating Image and Video Encoders for Enhanced Video Understanding

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T17:50:49.173346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:50:49.173346Z digest=sha256:41a61a6eb87b8ff00dc98c3284918a63a4bd1b0fa84390d306e53325a20ff278

Observation be38c7f1-2134-4f61-a717-636175509c2f · outbound

This paper cites Longvlm: Efficient long video understanding via large language models,.

When and What: Diffusion-Grounded VideoLLM with Entity Aware Segmentation for Long Video Understanding Longvlm: Efficient long video understanding via large language models,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:50:51.011850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T17:50:49.247131Z digest=sha256:5cc20ee82691c89e16bbdd1923d83f32e4f95eb46aaa4e6af9b1b18361cef004

Observation 627e59f1-c3ca-40f8-aa2d-712cb0a35502 · outbound

This paper cites Video summarization using denoising diffusion probabilistic model,.

When and What: Diffusion-Grounded VideoLLM with Entity Aware Segmentation for Long Video Understanding Video summarization using denoising diffusion probabilistic model,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:50:50.847614Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T17:50:49.390336Z digest=sha256:fc7444571e9ccd4943db44f33ce87f18e5796b0dda9f145555c7ecf507446a9b

Observation 9577ab86-3d90-4017-a53a-bb418e43127c · outbound

This paper cites Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos.

When and What: Diffusion-Grounded VideoLLM with Entity Aware Segmentation for Long Video Understanding Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T17:50:49.465440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:50:49.465440Z digest=sha256:a300238bd92394e98d9cd63e6716605d414767cff62a46750d9fe47a3550d1e8

Observation f9154289-0194-427f-acab-353318a02020 · outbound

This paper cites Videoglamm: A large multimodal model for pixel-level visual grounding in videos,.

When and What: Diffusion-Grounded VideoLLM with Entity Aware Segmentation for Long Video Understanding Videoglamm: A large multimodal model for pixel-level visual grounding in videos,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:50:50.709653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T17:50:49.612907Z digest=sha256:7439928fbbf207434c9408a0ad0af331f9e81ef754eaf38cecb619f00535f8b4

Observation e4090a78-d4c2-491d-ae9d-d464582bf38d · outbound

This paper cites VACE: All-in-One Video Creation and Editing.

When and What: Diffusion-Grounded VideoLLM with Entity Aware Segmentation for Long Video Understanding VACE: All-in-One Video Creation and Editing

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T17:50:49.717507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:50:49.717507Z digest=sha256:84a4b81924ebb0e2c4a1a901acf008cc4c32cfa92e024493ce6dc7f543875529

Observation 8ffbb9f7-fe06-45a1-8331-7f38b3832e13 · outbound

This paper cites Wan: Open and Advanced Large-Scale Video Generative Models.

When and What: Diffusion-Grounded VideoLLM with Entity Aware Segmentation for Long Video Understanding Wan: Open and Advanced Large-Scale Video Generative Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T17:50:49.870293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:50:49.870293Z digest=sha256:5201ab7bb6dee4867564a0bbe9cc14f3824ee176348f2ae82cf31ce87e892cc7

Observation ffe4b62a-7b5d-47bb-9c87-db8bf28ba84e · outbound

This paper cites Can i trust your answer? visually grounded video question answering,.

When and What: Diffusion-Grounded VideoLLM with Entity Aware Segmentation for Long Video Understanding Can i trust your answer? visually grounded video question answering,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:50:50.460962Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T17:50:50.011389Z digest=sha256:aa86ad5f92803034f867deccc4ca5635fd7bdece7fb5216c0a9b54b9af8355a3

Observation c5f7169c-70dc-4323-abb9-7951d3639879 · outbound

This paper cites Phi-4 Technical Report.

When and What: Diffusion-Grounded VideoLLM with Entity Aware Segmentation for Long Video Understanding Phi-4 Technical Report

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T17:50:50.085949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:50:50.085949Z digest=sha256:b6bca1a76ade9244d0a4d77b00499ecc889ebfdcd0cc769853999366bc074f76

Observation 26c83433-9103-4c3f-a97f-767d35310425 · outbound

This paper cites Grounded SAM: Assembling Open-World Models for Diverse Visual Tasks.

When and What: Diffusion-Grounded VideoLLM with Entity Aware Segmentation for Long Video Understanding Grounded SAM: Assembling Open-World Models for Diverse Visual Tasks

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T17:50:50.164729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:50:50.164729Z digest=sha256:dc8e5c77e6a5f9de0fbb7c0152f42480e3c4b5686eb95b677d609dbe01725bf8

Pith citing papers

Observation 5238bb53-07cb-4d26-88e8-8d378f7169a6 · inbound

YOSE: You Only Select Essential Tokens for Efficient DiT-based Video Object Removal cites this paper.

YOSE: You Only Select Essential Tokens for Efficient DiT-based Video Object Removal When and What: Diffusion-Grounded VideoLLM with Entity Aware Segmentation for Long Video Understanding

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:46:27.901967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-07T09:12:13.460169Z digest=sha256:8d104c5f3570b284f713508d8ddc086ff4643b5b92f2ffa211b5a7ae47cfb6fa