Pith. sign in

Paper Citation Record · LEDGER

Tarsier2: Advancing Large Vision-Language Models from Detailed Video Description to Comprehensive Video Understanding

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 27 inbound Pith citation observations for arXiv:2501.07888.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.07888 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 27 of 27 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 27 of 27 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T21:07:32.473373Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation ad0af566-b5c3-4249-9bc2-44d4c77eb671 · inbound

Goku: Flow Based Video Generative Foundation Models cites this paper.

Goku: Flow Based Video Generative Foundation Models Tarsier2: Advancing Large Vision-Language Models from Detailed Video Description to Comprehensive Video Understanding

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-08T21:07:32.473373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T21:07:32.473373Z digest=sha256:5e8916dd8ab3124f621a5baa974d07f931ff823e8fc34ec31e55d7ae25957655

Observation c4820ee3-cf4b-4940-b5c7-8710dff26a8d · inbound

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval cites this paper.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Tarsier2: Advancing Large Vision-Language Models from Detailed Video Description to Comprehensive Video Understanding

Reference 103

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:46.292799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:46.292799Z digest=sha256:9e75b2a5288b0e33a47e7b64572dd38addecabe0ba2bd20bf76fd85b8f5e83c2

Observation a745cefa-6f64-4205-8f2c-d09036d15a8b · inbound

Zooming from Context to Cue: Hierarchical Preference Optimization for Multi-Image MLLMs cites this paper.

Zooming from Context to Cue: Hierarchical Preference Optimization for Multi-Image MLLMs Tarsier2: Advancing Large Vision-Language Models from Detailed Video Description to Comprehensive Video Understanding

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T13:14:08.456890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:14:08.456890Z digest=sha256:c2639d8647655886c28ec388688a5b3932b718d46ad8723607b5de1cc30558cb

Observation 71e9c9c5-ee26-4ec7-af1a-7aa1e7509853 · inbound

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking cites this paper.

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking Tarsier2: Advancing Large Vision-Language Models from Detailed Video Description to Comprehensive Video Understanding

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:55.923555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:55.923555Z digest=sha256:46eeb0727d618907f22e0e0a3780d2b949c73cef83d5f4a30d18e0a2bf15cd2b

Observation c96d5b97-b387-450b-a5f1-13dc333d2a9f · inbound

How Important are Videos for Training Video LLMs? cites this paper.

How Important are Videos for Training Video LLMs? Tarsier2: Advancing Large Vision-Language Models from Detailed Video Description to Comprehensive Video Understanding

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T05:51:07.913992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:51:07.913992Z digest=sha256:a04faeaffca4a64fdd9aa99f8911ac1cae87ce241984e893b66f04bf70f161ec

Observation c4a72311-99ba-4367-80b0-3e7fe61effcf · inbound

Seedance 1.0: Exploring the Boundaries of Video Generation Models cites this paper.

Seedance 1.0: Exploring the Boundaries of Video Generation Models Tarsier2: Advancing Large Vision-Language Models from Detailed Video Description to Comprehensive Video Understanding

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-11T12:09:57.400029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-11T12:09:56.836351Z digest=sha256:30d5c8813d3dec5fda634b93906aafe0bf18b43d37faed7cc27a71aec76407af

Observation 506152f2-e699-4975-852f-f94d86af54a0 · inbound

V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning cites this paper.

V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning Tarsier2: Advancing Large Vision-Language Models from Detailed Video Description to Comprehensive Video Understanding

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:33:50.994322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-11T00:33:50.471804Z digest=sha256:1c2b2ad81e6aae0212c802eda644933bc213d8f33117de45de239df2868569bb

Observation a136efc8-662b-450c-90c5-22835e454eb3 · inbound

Kwai Keye-VL Technical Report cites this paper.

Kwai Keye-VL Technical Report Tarsier2: Advancing Large Vision-Language Models from Detailed Video Description to Comprehensive Video Understanding

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T20:45:10.093677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:45:10.093677Z digest=sha256:676b2fa30bc25da30c8788e218fe01ec1f12ea2d3459e9df9bdeceafc45391bc

Observation 5672ce2c-f641-4552-ba1f-88334ca51faf · inbound

Kwai Keye-VL 1.5 Technical Report cites this paper.

Kwai Keye-VL 1.5 Technical Report Tarsier2: Advancing Large Vision-Language Models from Detailed Video Description to Comprehensive Video Understanding

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-05T12:28:28.452005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:28:28.452005Z digest=sha256:232b666325f21e1094ee6b5771cf29f704683126acf72fcadf9658aba0ef1b0e

Observation 71939508-e1c3-4376-8bb6-6eb0b7b8a73a · inbound

NeMo: Needle in a Montage for Video-Language Understanding cites this paper.

NeMo: Needle in a Montage for Video-Language Understanding Tarsier2: Advancing Large Vision-Language Models from Detailed Video Description to Comprehensive Video Understanding

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-04T13:54:25.161492Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T13:54:25.161492Z digest=sha256:33bd9b751fd08b8fb49e1f63dbdaf46d2cdd965ace52f5d5d0c26eaed02f59cb

Observation d8945482-6b3c-4c05-8c01-c345ae36d1ed · inbound

JoVA: Unified Multimodal Learning for Joint Video-Audio Generation and Editing cites this paper.

JoVA: Unified Multimodal Learning for Joint Video-Audio Generation and Editing Tarsier2: Advancing Large Vision-Language Models from Detailed Video Description to Comprehensive Video Understanding

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-03T16:25:59.808955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:25:59.808955Z digest=sha256:b5ef553f9c8c656867b695823e4be0c3d1fa4ad7295fa7d08db75f2534a923f1

Observation 837614d4-94f0-44a2-b888-7a5e2b1dad60 · inbound

Syn4D: A Multiview Synthetic 4D Dataset cites this paper.

Syn4D: A Multiview Synthetic 4D Dataset Tarsier2: Advancing Large Vision-Language Models from Detailed Video Description to Comprehensive Video Understanding

Reference 130

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T17:46:07.659960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-08T17:14:50.827455Z digest=sha256:e2c6f0f0ae9281fbc11d0645667c5c0091366c0673861d0f21186e468e612f6c

Observation 97f91e0e-d866-4983-9f5d-63cfd6aafa82 · inbound

Syn4D: A Multiview Synthetic 4D Dataset cites this paper.

Syn4D: A Multiview Synthetic 4D Dataset Tarsier2: Advancing Large Vision-Language Models from Detailed Video Description to Comprehensive Video Understanding

Reference 129

Resolution
unresolved
no resolver link, observed 2026-07-12T17:33:21.224727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-12T17:33:21.224727Z digest=sha256:58ce5ac393171a00bd7e95c687775e05885bdab7ffca5c4e5865a5b8b2416798

Observation 81f5c3bc-281a-4040-813c-dd57eb4f54a0 · inbound

TOC-Bench: A Temporal Object Consistency Benchmark for Video Large Language Models cites this paper.

TOC-Bench: A Temporal Object Consistency Benchmark for Video Large Language Models Tarsier2: Advancing Large Vision-Language Models from Detailed Video Description to Comprehensive Video Understanding

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:31:26.810795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-12T04:13:21.487431Z digest=sha256:e02a7b5e713970af98393ae096bcb3cab8865866e2d811a46806d7b9cfd5f781

Observation 2fee3852-83c5-41a1-9c57-bdca37c063c4 · inbound

TOC-Bench: A Temporal Object Consistency Benchmark for Video Large Language Models cites this paper.

TOC-Bench: A Temporal Object Consistency Benchmark for Video Large Language Models Tarsier2: Advancing Large Vision-Language Models from Detailed Video Description to Comprehensive Video Understanding

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-13T06:57:28.247244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T06:53:42.726350Z digest=sha256:aff3f100cc1faa686d4f9e824ac031722586324d1aef774e881ae05cb4d0b4de

Observation c13f9945-5d2d-4cd3-b539-cdc63e8a7fcd · inbound

GazeBehavior Annotation Toolkit (GBAT): AI-powered toolkit for automatic annotation of egocentric eye-tracking and video data of child-caregiver interaction cites this paper.

GazeBehavior Annotation Toolkit (GBAT): AI-powered toolkit for automatic annotation of egocentric eye-tracking and video data of child-caregiver interaction Tarsier2: Advancing Large Vision-Language Models from Detailed Video Description to Comprehensive Video Understanding

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-25T05:50:24.582603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-25T05:46:57.253756Z digest=sha256:f46e8f3f63cbd87cf983ae3b762933c4e7eb7f3ee7a999e6a4b0fc1fb20be4b6

Observation e825296b-f6c7-4712-8e04-fc4d80848624 · inbound

LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence cites this paper.

LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence Tarsier2: Advancing Large Vision-Language Models from Detailed Video Description to Comprehensive Video Understanding

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-06-29T22:13:59.509917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T22:12:05.365596Z digest=sha256:468f3b4b8c1b63ce5fe093da1c5fbe8bc1fa87675a06ef95ed7632c2361a34ff

Observation 2e656377-e5d8-4dd9-a5dd-4f8f19bbae3b · inbound

VCap: Hypergeometric Rewards for Weak-to-Strong Visual Captioning cites this paper.

VCap: Hypergeometric Rewards for Weak-to-Strong Visual Captioning Tarsier2: Advancing Large Vision-Language Models from Detailed Video Description to Comprehensive Video Understanding

Reference 61

Resolution
verified exact
arxiv_id, observed 2026-06-29T13:23:28.467011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T13:13:57.599970Z digest=sha256:39a604fedbb0fe56b7fe157043f6835fc64a5f8114b9676a537d25162f7b2d2d

Observation 2d08a864-9ca4-44ec-9909-28ed378eedee · inbound

Task-Focused Memorization for Multimodal Agents cites this paper.

Task-Focused Memorization for Multimodal Agents Tarsier2: Advancing Large Vision-Language Models from Detailed Video Description to Comprehensive Video Understanding

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-06-29T00:12:50.491227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T23:14:45.287255Z digest=sha256:3d6bab714288b0e96a8f5bb1105c176f6bf519367c1e2f60025a969e09d9478c

Observation d344e5c1-beb9-4a5f-a9ab-a1c298b9f080 · inbound

Watch, Remember, Reason: Human-View Video Understanding with MLLMs cites this paper.

Watch, Remember, Reason: Human-View Video Understanding with MLLMs Tarsier2: Advancing Large Vision-Language Models from Detailed Video Description to Comprehensive Video Understanding

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:27:15.563714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T22:00:28.350003Z digest=sha256:06303c0d0ae8e324406de4c5adb7c86fb2d7ddb552e3eb933a223f3f681a7749

Observation 8cce8127-212d-43fe-8fed-4edf35c38d80 · inbound

CapRL++: Unified Reinforcement Learning with Verifiable Rewards for Dense Image and Video Captioning cites this paper.

CapRL++: Unified Reinforcement Learning with Verifiable Rewards for Dense Image and Video Captioning Tarsier2: Advancing Large Vision-Language Models from Detailed Video Description to Comprehensive Video Understanding

Reference 95

Resolution
verified exact
arxiv_id, observed 2026-07-03T00:17:29.003700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T17:21:38.543724Z digest=sha256:ae10b3f3a9f311d3fe8ca913173ea175aae55135589098ae95e55bcc2dbfdace

Observation 28359097-2996-405b-9bfd-b14c2bf3938d · inbound

CineCap: Structured Reasoning with Spatio-Temporal Anchors for Cinematographic Video Captioning cites this paper.

CineCap: Structured Reasoning with Spatio-Temporal Anchors for Cinematographic Video Captioning Tarsier2: Advancing Large Vision-Language Models from Detailed Video Description to Comprehensive Video Understanding

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-07-04T17:40:00.898737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-25T23:29:24.520537Z digest=sha256:86b9013da95024c59f12f2d21c094c0a44a73cca530f01e18e9ed4ae39e3887c

Observation d43b88f3-90dd-4d8c-b055-ba896156cbef · inbound

MotionAtlas: Detailed Region Captioning for Motion-Centric Videos cites this paper.

MotionAtlas: Detailed Region Captioning for Motion-Centric Videos Tarsier2: Advancing Large Vision-Language Models from Detailed Video Description to Comprehensive Video Understanding

Reference 43

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T07:24:21.165380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T07:21:46.783970Z digest=sha256:e65e154d3190c7b0e940a51f3965dd7f4ca47668429b50266a0c2b69130a96b3

Observation bba6e8b5-3184-4c67-af0b-ba8051221a73 · inbound

MoHallBench: A Benchmark for Motion Hallucination in Video Large Language Models cites this paper.

MoHallBench: A Benchmark for Motion Hallucination in Video Large Language Models Tarsier2: Advancing Large Vision-Language Models from Detailed Video Description to Comprehensive Video Understanding

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-07-02T13:46:58.616375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-02T13:44:48.250838Z digest=sha256:7dff2520c87737ba09653a8e90e0b1fab20e8bf9fd166d4c675efcf6c4c4ebfa

Observation e23db478-1672-4805-8959-7441d03e82d8 · inbound

Temporal and Cross-Modal Alignment for Enhanced Audiovisual Video Captioning cites this paper.

Temporal and Cross-Modal Alignment for Enhanced Audiovisual Video Captioning Tarsier2: Advancing Large Vision-Language Models from Detailed Video Description to Comprehensive Video Understanding

Reference 66

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T16:38:39.648951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-03T16:37:06.384435Z digest=sha256:6de743481fe899045517220232ade75134efbbdf20adc35a1f2c62b647cf8a14

Observation dc8bdbe5-db5f-4970-bc1b-09f5f5777cae · inbound

VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding cites this paper.

VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Tarsier2: Advancing Large Vision-Language Models from Detailed Video Description to Comprehensive Video Understanding

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-02T00:44:43.581545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:44:43.581545Z digest=sha256:77dddb45fd6fadcba9f48d350cb4077e0fd1675c701236027e75112387a8c112

Observation 5a5680dd-462b-49e5-a284-ee7adb09ce1f · inbound

LookME: Lookup-Based Multimodal Embeddings for Layer Injection in Vision-Language Models cites this paper.

LookME: Lookup-Based Multimodal Embeddings for Layer Injection in Vision-Language Models Tarsier2: Advancing Large Vision-Language Models from Detailed Video Description to Comprehensive Video Understanding

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-02T06:33:14.955758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:33:14.955758Z digest=sha256:4aa48d7ad540f8ca6a2d48df6fe980137b28aef8222fef3f99fedffee8e40b29