Pith. sign in

Paper Citation Record · LEDGER

Dispider: Enabling Video LLMs with Active Real-Time Interaction via Disentangled Perception, Decision, and Reaction

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 12 inbound Pith citation observations for arXiv:2501.03218.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.03218 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 12 of 12 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 12 of 12 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T20:06:35.167362Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T22:26:17.926283Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 8a52dff9-a257-442d-8a57-e889f9000459 · inbound

VideoRoPE: What Makes for Good Video Rotary Position Embedding? cites this paper.

VideoRoPE: What Makes for Good Video Rotary Position Embedding? Dispider: Enabling Video LLMs with Active Real-Time Interaction via Disentangled Perception, Decision, and Reaction

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-08T20:06:35.167362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:06:35.167362Z digest=sha256:394e68a4b94dd9ee8f97be8bd2f92831a745535aece29b3c615e51db607355c4

Observation 88a2eff5-a20f-4274-bfa0-802292c48365 · inbound

LiveVLM: Efficient Online Video Understanding via Streaming-Oriented KV Cache and Retrieval cites this paper.

LiveVLM: Efficient Online Video Understanding via Streaming-Oriented KV Cache and Retrieval Dispider: Enabling Video LLMs with Active Real-Time Interaction via Disentangled Perception, Decision, and Reaction

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-22T14:31:40.751212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-22T14:26:59.015559Z digest=sha256:5cab91eb5d0fba3bfb098470d4198a12e98f5f916cdb6b4cce9fa5da482fd056

Observation 7b0e7426-b2f9-414d-8ade-fed24f1340f5 · inbound

ProactiveVideoQA: A Comprehensive Benchmark Evaluating Proactive Interactions in Video Large Language Models cites this paper.

ProactiveVideoQA: A Comprehensive Benchmark Evaluating Proactive Interactions in Video Large Language Models Dispider: Enabling Video LLMs with Active Real-Time Interaction via Disentangled Perception, Decision, and Reaction

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T18:03:14.023334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:03:14.023334Z digest=sha256:9753544447e5c44e24bcc58a4ef6ab78d9954023f87f7cf544989862d8dd00f7

Observation 53cf1a43-0fee-4632-907d-776919048db0 · inbound

HumanPCR: Probing MLLM Capabilities in Diverse Human-Centric Scenes cites this paper.

HumanPCR: Probing MLLM Capabilities in Diverse Human-Centric Scenes Dispider: Enabling Video LLMs with Active Real-Time Interaction via Disentangled Perception, Decision, and Reaction

Reference 104

Resolution
unresolved
no resolver link, observed 2026-08-05T19:03:11.732167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:03:11.732167Z digest=sha256:bce26cfc0469dc5cd467d4bd48d98d424ec3c204c6719661e3a6ceb7456bdb72

Observation 95f770cf-580c-4219-85d1-db4bf6c3baea · inbound

Boosting Embodied AI Agents through Perception-Generation Disaggregation and Asynchronous Pipeline Execution cites this paper.

Boosting Embodied AI Agents through Perception-Generation Disaggregation and Asynchronous Pipeline Execution Dispider: Enabling Video LLMs with Active Real-Time Interaction via Disentangled Perception, Decision, and Reaction

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-04T18:58:12.868093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:58:12.868093Z digest=sha256:762303dbe6d72d28cc61f7dc28b450472f720763e5dc110c3fa2c8b2516a265c

Observation 40074e99-5415-4b8a-9da5-fb57567a309d · inbound

Streaming Video Instruction Tuning cites this paper.

Streaming Video Instruction Tuning Dispider: Enabling Video LLMs with Active Real-Time Interaction via Disentangled Perception, Decision, and Reaction

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-16T19:48:21.763487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T19:44:11.032898Z digest=sha256:569b99459f0d036554ea5ab49e6a22e4cba6ac0bcbba09d1097d09a727a7e9e9

Observation 6362aea4-172e-4120-8ee6-f36894c11b30 · inbound

HERMES: KV Cache as Hierarchical Memory for Efficient Streaming Video Understanding cites this paper.

HERMES: KV Cache as Hierarchical Memory for Efficient Streaming Video Understanding Dispider: Enabling Video LLMs with Active Real-Time Interaction via Disentangled Perception, Decision, and Reaction

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-16T12:57:53.912952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T12:55:04.564442Z digest=sha256:af6d84b3933864d17026dc7ef0844be61ae0cc7455b546a15ba2de73b8e9e707

Observation e46447b3-5db4-4e6e-933b-6df5e1f0f461 · inbound

SVCBench: A Streaming Video Counting Benchmark for Spatial-Temporal State Maintenance cites this paper.

SVCBench: A Streaming Video Counting Benchmark for Spatial-Temporal State Maintenance Dispider: Enabling Video LLMs with Active Real-Time Interaction via Disentangled Perception, Decision, and Reaction

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-14T22:09:56.270516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T22:09:56.270516Z digest=sha256:f1150b3dcc57b7534201a410da9a82a2b6eec8955ec76946772e324829b324ee

Observation 16ce67ef-091d-4468-a7a9-0547ad079b39 · inbound

MOSS-Video-Preview: Toward Real-Time Video Understanding via Cross-Attention cites this paper.

MOSS-Video-Preview: Toward Real-Time Video Understanding via Cross-Attention Dispider: Enabling Video LLMs with Active Real-Time Interaction via Disentangled Perception, Decision, and Reaction

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-07-01T22:26:17.927733Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T15:22:31.310003Z digest=sha256:bbbc826a3fe9e274bd57f720c0182bf1bad68720135d84a10f6f8245f85a7126

Observation dc3c6d0c-8ff1-470e-a201-2abe63c3863e · inbound

Agent-Computer Observation Interfaces Enable Dynamic Computer Use cites this paper.

Agent-Computer Observation Interfaces Enable Dynamic Computer Use Dispider: Enabling Video LLMs with Active Real-Time Interaction via Disentangled Perception, Decision, and Reaction

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-06-30T07:04:21.262807Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T06:59:20.295818Z digest=sha256:8120f7699402535877a62ed1a31538c298c53e269a3326d6d9f2dbd15979c9aa

Observation bc06c397-3e9e-4ed2-b1e6-8c1814e68646 · inbound

Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model cites this paper.

Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model Dispider: Enabling Video LLMs with Active Real-Time Interaction via Disentangled Perception, Decision, and Reaction

Reference 71

Resolution
unresolved
no resolver link, observed 2026-07-31T06:20:13.871291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:20:13.871291Z digest=sha256:89be0082ac23d678a39be3bb92bb85e8abc7a04eb15eb4d4fa93b81c1522990c

Observation d64cfaea-a437-48a9-9970-1ef6eed97e19 · inbound

ObjectStream: Latent Objects as Memory Anchors for Streaming Video Understanding cites this paper.

ObjectStream: Latent Objects as Memory Anchors for Streaming Video Understanding Dispider: Enabling Video LLMs with Active Real-Time Interaction via Disentangled Perception, Decision, and Reaction

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-04T03:21:46.866199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T03:21:46.866199Z digest=sha256:4ad8bde097e0a7420e5e6503a9187278396e7b410a268076d65c922c18b6bf5c