Pith. sign in

Paper Citation Record · LEDGER

PixFoundation 2.0: Do Video Multi-Modal LLMs Use Motion in Visual Grounding?

As of 16 August 2026, this Paper Citation Record lists 18 of 18 outbound references and 0 inbound Pith citation observations for arXiv:2509.02807.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.02807 v1

Coverage vector

measured 18 of 18 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T11:27:47.145067Z

measured 18 of 18 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

18 of 18 outbound references displayed

  • verified exact1
  • verified fuzzy1
  • unresolved16
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4018e7bd-4846-4503-be68-928453b8ea2f · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

PixFoundation 2.0: Do Video Multi-Modal LLMs Use Motion in Visual Grounding? Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T11:27:47.077290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:27:47.077290Z digest=sha256:6373d143dae3ce7e11d1fc69688e2e46eefd0a55f156e8ae7b2e30994478f13f

Observation 2c9655e2-ebaa-4cb8-84bb-095876763e1c · outbound

This paper cites MED-VT++: Unifying Multimodal Learning with a Multiscale Encoder-Decoder Video Transformer.

PixFoundation 2.0: Do Video Multi-Modal LLMs Use Motion in Visual Grounding? MED-VT++: Unifying Multimodal Learning with a Multiscale Encoder-Decoder Video Transformer

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-08-05T11:27:47.523163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-05T11:27:47.082104Z digest=sha256:bd4ad60120ede462272a4c13fe6dc79ec6ded6a57982d4608364ae466e047690

Observation 004ce299-6190-41c0-9871-e2cded83df37 · outbound

This paper cites Visual instruction tuning.

PixFoundation 2.0: Do Video Multi-Modal LLMs Use Motion in Visual Grounding? Visual instruction tuning

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T11:27:47.602290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-05T11:27:47.092352Z digest=sha256:02c98dcf962c4197cfaa1bcabbe4fda4acba8292800265ae3541b77623b87982

Observation eba6ded5-0332-4278-a22e-9cfa58f8dd31 · outbound

This paper cites Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models.

PixFoundation 2.0: Do Video Multi-Modal LLMs Use Motion in Visual Grounding? Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T11:27:47.097393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:27:47.097393Z digest=sha256:ed7c32a3314a8382ec205daa74a353a90cc522ecfd755194816a51c0baa03b53

Observation cbaa8710-72bb-4e3c-b023-83dabb5b3358 · outbound

This paper cites The 2017 DAVIS Challenge on Video Object Segmentation.

PixFoundation 2.0: Do Video Multi-Modal LLMs Use Motion in Visual Grounding? The 2017 DAVIS Challenge on Video Object Segmentation

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T11:27:47.106806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:27:47.106806Z digest=sha256:b8448e7e10ff8b6574dfa19044e7fdcafeed854234a5088714412f5c9d952fac

Observation b6705038-e826-46a5-93ff-615f14bea5d4 · outbound

This paper cites SAM 2: Segment Anything in Images and Videos.

PixFoundation 2.0: Do Video Multi-Modal LLMs Use Motion in Visual Grounding? SAM 2: Segment Anything in Images and Videos

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T11:27:47.111735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:27:47.111735Z digest=sha256:0d46e062768eb97bbc0927960219f4c1aa7fea6afe767a98a5f9f119747c61e3

Observation 85c6c5a5-706f-4501-b376-a90ea40fade8 · outbound

This paper cites Seeing the arrow of time in large multimodal models.

PixFoundation 2.0: Do Video Multi-Modal LLMs Use Motion in Visual Grounding? Seeing the arrow of time in large multimodal models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T11:27:47.126396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:27:47.126396Z digest=sha256:5e6e37e8a5e0866b7cf829d8d9a4970fbd0db091e02f88549e5db3831a6abc71

Observation 06762684-0978-4228-8d19-64855a174207 · outbound

This paper cites OMG-LLaVA: Bridging Image-level, Object-level, Pixel-level Reasoning and Understanding.

PixFoundation 2.0: Do Video Multi-Modal LLMs Use Motion in Visual Grounding? OMG-LLaVA: Bridging Image-level, Object-level, Pixel-level Reasoning and Understanding

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T11:27:47.135264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:27:47.135264Z digest=sha256:0fea439cf40c3ab4959b5e091067cdfe54f95519321d59340c1d840cdb1fbeaa

Observation 48401b11-edf3-420b-9e40-32acd7052281 · outbound

This paper cites InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models.

PixFoundation 2.0: Do Video Multi-Modal LLMs Use Motion in Visual Grounding? InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T11:27:47.140036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:27:47.140036Z digest=sha256:3d8a87bf23319ae10d36c788658ffeccc615730c44ada8fcf6620d1e773dc28d

Observation 4ebc1d82-7cb1-49f5-9264-9aaf78b8426d · outbound

This paper cites Apollo: An Exploration of Video Understanding in Large Multimodal Models.

PixFoundation 2.0: Do Video Multi-Modal LLMs Use Motion in Visual Grounding? Apollo: An Exploration of Video Understanding in Large Multimodal Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T11:27:47.145067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:27:47.145067Z digest=sha256:2c99ae8afd173eecd4d15a8c29c8010131d411db4d895762bf750d4f6a91bc14

Observation 6e33aa3b-2862-416a-ab2a-c549c15742d1 · outbound

This paper cites Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos.

PixFoundation 2.0: Do Video Multi-Modal LLMs Use Motion in Visual Grounding? Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-05T11:27:47.130766Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:27:47.130766Z digest=sha256:6745e5aa4ee80d43071fb5f287d4ba25fc21973b077af8f67912986f42add963

Observation 95e3f717-51cf-4246-8c63-32409f313d6f · outbound

This paper cites Video-LLaVA: Learning United Visual Representation by Alignment Before Projection.

PixFoundation 2.0: Do Video Multi-Modal LLMs Use Motion in Visual Grounding? Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-05T11:27:47.086884Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:27:47.086884Z digest=sha256:e2c254ed61e94009cb047067a28bbe2d7d76e61654b5bb18ec449dd5feeaa83e

Observation 517b939d-d764-4c18-9746-ecb1f30d78b0 · outbound

This paper cites Pixfoundation: Are we heading in the right direction with pixel-level vision foundation models? arXiv preprint arXiv:2502.04192,.

PixFoundation 2.0: Do Video Multi-Modal LLMs Use Motion in Visual Grounding? Pixfoundation: Are we heading in the right direction with pixel-level vision foundation models? arXiv preprint arXiv:2502.04192,

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-05T11:27:47.117264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:27:47.117264Z digest=sha256:0c2398da873ec2d9722ff565205e16977d394c7112209af048104ed2287a40ce

Observation f13312fa-4a4e-4185-93fd-866c791bac9f · outbound

This paper cites VideoGLaMM: A Large Multimodal Model for Pixel-Level Visual Grounding in Videos.

PixFoundation 2.0: Do Video Multi-Modal LLMs Use Motion in Visual Grounding? VideoGLaMM: A Large Multimodal Model for Pixel-Level Visual Grounding in Videos

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-05T11:27:47.102050Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:27:47.102050Z digest=sha256:00816cb7d4b9aeb9e388b147d46d362c7ed29aa69a2cc512e1a49b4e0377e9f2

Observation bc7b336e-7c28-4247-84c1-c97ccebf12f9 · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

PixFoundation 2.0: Do Video Multi-Modal LLMs Use Motion in Visual Grounding? Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-05T11:27:47.072265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:27:47.072265Z digest=sha256:bfeedf83d539f55989a5a6b162b04f1b67cea56e1d68f75aad7bae7f0301f857

Observation 21ce8e8d-0df4-4b25-bdb4-0b5dcea8cd19 · outbound

This paper cites Qwen2.5-VL Technical Report.

PixFoundation 2.0: Do Video Multi-Modal LLMs Use Motion in Visual Grounding? Qwen2.5-VL Technical Report

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-05T11:27:47.067394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:27:47.067394Z digest=sha256:f7cf84a9d01196c25b85a0a8d476eadc76a9c184a7f024d3ac31c3b60ffde077

Observation 0f7d0fa9-84e4-4575-b4cc-093f85176bcc · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

PixFoundation 2.0: Do Video Multi-Modal LLMs Use Motion in Visual Grounding? Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-05T11:27:47.062248Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:27:47.062248Z digest=sha256:4f99ba666c5727c14e0134855aae8e62b171a0cb26dccc9474586c9cf6c5258e

Observation 6bafd136-ad1a-413c-96a7-1c0255ab1620 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

PixFoundation 2.0: Do Video Multi-Modal LLMs Use Motion in Visual Grounding? Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-05T11:27:47.121911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:27:47.121911Z digest=sha256:b6e7841b3e757712f48766d10b50fdbe609c9f30306dbfb46504f7068ffce0ea

Pith citing papers

No inbound Pith citation observations are available.