Pith. sign in

Paper Citation Record · LEDGER

LLaVA-ST: A Multimodal Large Language Model for Fine-Grained Spatial-Temporal Understanding

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 7 inbound Pith citation observations for arXiv:2501.08282.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.08282 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 7 of 7 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:35:45.865632Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-19T13:17:18.523647Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 8994b48c-70a8-44b5-8d48-9514de05ad00 · inbound

MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment Grounding cites this paper.

MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment Grounding LLaVA-ST: A Multimodal Large Language Model for Fine-Grained Spatial-Temporal Understanding

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-19T13:17:18.526285Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-19T13:13:40.485342Z digest=sha256:00d013242d5a4e9e8d8c68f196d00d36ae5cdc1ffbd3d4c566caf93c57e6b6d0

Observation f79c5a1c-78e4-4915-8011-72b2ac8f5279 · inbound

Reinforcement Learning Tuning for VideoLLMs: Reward Design and Data Efficiency cites this paper.

Reinforcement Learning Tuning for VideoLLMs: Reward Design and Data Efficiency LLaVA-ST: A Multimodal Large Language Model for Fine-Grained Spatial-Temporal Understanding

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T11:35:45.865632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:35:45.865632Z digest=sha256:ada416a27cd5b516f1eab4d7068e51c5cb20522502991023c177d77238261d35

Observation 396438cd-7478-4801-ac34-1f5cec8d7300 · inbound

ComRoPE: Scalable and Robust Rotary Position Embedding Parameterized by Trainable Commuting Angle Matrices cites this paper.

ComRoPE: Scalable and Robust Rotary Position Embedding Parameterized by Trainable Commuting Angle Matrices LLaVA-ST: A Multimodal Large Language Model for Fine-Grained Spatial-Temporal Understanding

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T11:03:17.170264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:03:17.170264Z digest=sha256:88366f1c8826356b792beb43cb02083108507680cc111c817bb7773acbada743

Observation 703cec36-f077-4afb-9766-b3706ce35412 · inbound

A Survey on Video Temporal Grounding with Multimodal Large Language Model cites this paper.

A Survey on Video Temporal Grounding with Multimodal Large Language Model LLaVA-ST: A Multimodal Large Language Model for Fine-Grained Spatial-Temporal Understanding

Reference 114

Resolution
unresolved
no resolver link, observed 2026-08-05T23:32:17.977612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:32:17.977612Z digest=sha256:0a84b02f65c3288fbca004fc63543ae49914cebcbb0a01233076fc55d934eef5

Observation 323ef9e9-47cc-41dc-8691-b32fc052179e · inbound

STER-VLM: Spatio-Temporal With Enhanced Reference Vision-Language Models cites this paper.

STER-VLM: Spatio-Temporal With Enhanced Reference Vision-Language Models LLaVA-ST: A Multimodal Large Language Model for Fine-Grained Spatial-Temporal Understanding

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T19:06:23.616871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:06:23.616871Z digest=sha256:cda5e77def86984dbcc75ce0e62b921bafd9f853c9d8c0f33f78fc110c7ee89e

Observation 2290a2b5-f67c-4115-89b0-d84a337444eb · inbound

AeroDuo: Aerial Duo for UAV-based Vision and Language Navigation cites this paper.

AeroDuo: Aerial Duo for UAV-based Vision and Language Navigation LLaVA-ST: A Multimodal Large Language Model for Fine-Grained Spatial-Temporal Understanding

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-05T18:06:00.350739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T18:06:00.350739Z digest=sha256:77b608507fb3dcc901fba50a373a8b8b392ac9d528b1ee704817636cbe4d264e

Observation 3fb88a8f-e1f4-4ad4-aef7-ba939d81714e · inbound

ViLL-E: Video LLM Embeddings for Retrieval cites this paper.

ViLL-E: Video LLM Embeddings for Retrieval LLaVA-ST: A Multimodal Large Language Model for Fine-Grained Spatial-Temporal Understanding

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-11T11:21:01.954274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T15:00:43.573409Z digest=sha256:a7553e6754324362c62e53c61a60245da4e38068928b4eac8c8c78852b5cdee6