Pith. sign in

Paper Citation Record · LEDGER

ViLLa: Video Reasoning Segmentation with Large Language Model

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 15 inbound Pith citation observations for arXiv:2407.14500.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2407.14500 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 15 of 15 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 15 of 15 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:27:19.300329Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T15:09:55.216469Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 59d82159-fd8b-4c8e-b1ef-e006c7f06415 · inbound

Reasoning Segmentation for Images and Videos: A Survey cites this paper.

Reasoning Segmentation for Images and Videos: A Survey ViLLa: Video Reasoning Segmentation with Large Language Model

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:19.300329Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:19.300329Z digest=sha256:402832c910e5b86ae8af5b9961d22c5e73c3a0a3e53aabcf43088140f8a91666

Observation 51c3e5c4-7511-4120-a3b2-65b765c8c9b2 · inbound

InterRVOS: Interaction-aware Referring Video Object Segmentation cites this paper.

InterRVOS: Interaction-aware Referring Video Object Segmentation ViLLa: Video Reasoning Segmentation with Large Language Model

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T11:29:29.149297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:29:29.149297Z digest=sha256:d90c09a696a95a932fe761520f8f1310e1c7fda7c134901d344da323a00ab912

Observation ab1af148-e5f5-45e1-86d7-1042694e9af6 · inbound

A Comprehensive Survey on Video Scene Parsing:Advances, Challenges, and Prospects cites this paper.

A Comprehensive Survey on Video Scene Parsing:Advances, Challenges, and Prospects ViLLa: Video Reasoning Segmentation with Large Language Model

Reference 200

Resolution
unresolved
no resolver link, observed 2026-08-07T00:34:25.104905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:34:25.104905Z digest=sha256:3ab6ffb4e068de7d0656ef8330c0795a0a886b2da8da115236294d9cf493f2c7

Observation d69fd75a-9aa7-4bd6-928c-b428aaaf0e05 · inbound

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs cites this paper.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs ViLLa: Video Reasoning Segmentation with Large Language Model

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:20.326601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:20.326601Z digest=sha256:08420de81ba7f8086bf1dde19df11e86c7e4b1d0ffee6389473dc825d5b21a2b

Observation 51c74104-c7d0-4941-bc6f-b74ebc17b510 · inbound

HRSeg: High-Resolution Visual Perception and Enhancement for Reasoning Segmentation cites this paper.

HRSeg: High-Resolution Visual Perception and Enhancement for Reasoning Segmentation ViLLa: Video Reasoning Segmentation with Large Language Model

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-06T16:43:56.438861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:43:56.438861Z digest=sha256:740171625371b51b1d2dedd91b1c1f2428aaa635dc8f5b973579608a46710626

Observation f35ef5b4-7b2b-4406-9312-0d74b021a834 · inbound

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation cites this paper.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation ViLLa: Video Reasoning Segmentation with Large Language Model

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-06T11:16:10.460552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:16:10.460552Z digest=sha256:31cfb040ba876d4e91e05a12c1f83cc4ebfc33931b6f374af0d5bcc888d5d00a

Observation dd204a70-d312-42b5-b8a4-8b8f3464d726 · inbound

Unleashing Hierarchical Reasoning: An LLM-Driven Framework for Training-Free Referring Video Object Segmentation cites this paper.

Unleashing Hierarchical Reasoning: An LLM-Driven Framework for Training-Free Referring Video Object Segmentation ViLLa: Video Reasoning Segmentation with Large Language Model

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-05T05:09:31.781108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T05:09:31.781108Z digest=sha256:4042cc83a1f16055e9a6b395853e0cb83c3edeaa4c5145ecab7ab718cb12436b

Observation b829bba1-a116-4cf6-a6c0-05e336f5e66c · inbound

LMMs Meet Object-Centric Vision: Understanding, Segmentation, Editing and Generation cites this paper.

LMMs Meet Object-Centric Vision: Understanding, Segmentation, Editing and Generation ViLLa: Video Reasoning Segmentation with Large Language Model

Reference 239

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:11:08.070769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T15:35:37.095627Z digest=sha256:92b7d44c21123d7c048a2383d4c1a9386bf0c366e45ee1eabd9bb6da667272c8

Observation 9752265a-1874-4753-a54e-5512d78a2de7 · inbound

PixDLM: A Dual-Path Multimodal Language Model for UAV Reasoning Segmentation cites this paper.

PixDLM: A Dual-Path Multimodal Language Model for UAV Reasoning Segmentation ViLLa: Video Reasoning Segmentation with Large Language Model

Reference 61

Resolution
verified exact
arxiv_id, observed 2026-05-10T09:23:37.143330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T09:20:54.635375Z digest=sha256:ec14f04076cf984d3afef175fe09d7dfb08939b12f2ab59c2fdea47541c53438

Observation 41eec131-2845-444a-84fa-8165ba0eec00 · inbound

PixDLM: A Dual-Path Multimodal Language Model for UAV Reasoning Segmentation cites this paper.

PixDLM: A Dual-Path Multimodal Language Model for UAV Reasoning Segmentation ViLLa: Video Reasoning Segmentation with Large Language Model

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-02T16:10:02.936737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T16:10:02.936737Z digest=sha256:5167a39ef55311cd947fa6422b0166655cb8a98ee14ecd2e086f0bc0fdc8de1c

Observation 24c37d64-4d13-4d55-bf73-140cbe1b3500 · inbound

Weakly-Supervised Referring Video Object Segmentation through Text Supervision cites this paper.

Weakly-Supervised Referring Video Object Segmentation through Text Supervision ViLLa: Video Reasoning Segmentation with Large Language Model

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-10T05:56:11.089619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T05:54:33.227975Z digest=sha256:e86151501df345235a1a159e458fd45a7b674b20e1df2e7f7a6bb4312bf4d554

Observation 88f36ba8-30bd-4fa6-9265-6d15d5212724 · inbound

APRVOS: 1st Place Winner of 5th PVUW MeViS-Audio Track cites this paper.

APRVOS: 1st Place Winner of 5th PVUW MeViS-Audio Track ViLLa: Video Reasoning Segmentation with Large Language Model

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-10T03:29:21.565364Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T03:29:12.783993Z digest=sha256:10ecd54198a469dd9a7b71c1b6698a0681a31dd23a728ad8d9f90a5b90055d98

Observation 3b99a5eb-097e-40f2-838c-751e1f2a28ae · inbound

AgentRVOS for MeViS-Text Track of 5th PVUW Challenge: 3rd Method cites this paper.

AgentRVOS for MeViS-Text Track of 5th PVUW Challenge: 3rd Method ViLLa: Video Reasoning Segmentation with Large Language Model

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-10T09:33:41.857762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T05:14:25.423302Z digest=sha256:e1ba7847f04019c499b0dd890ccb97e5f10c937ce989826860b6312045822295

Observation 6e2118be-67cd-429d-b53a-d9e35271a891 · inbound

RCoT-Seg: Reinforced Chain-of-Thought for Video Reasoning and Segmentation cites this paper.

RCoT-Seg: Reinforced Chain-of-Thought for Video Reasoning and Segmentation ViLLa: Video Reasoning Segmentation with Large Language Model

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:30:59.882045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-11T01:16:25.031349Z digest=sha256:d9f31762aea7d27dd73a72fab6c4c9b0d4191b9dc6c0b58b89a87c80e1e81a8e

Observation 7c45d39c-2791-40cc-8f92-915504edf8e5 · inbound

From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models cites this paper.

From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models ViLLa: Video Reasoning Segmentation with Large Language Model

Reference 84

Resolution
verified exact
arxiv_id, observed 2026-07-04T15:09:55.217966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-26T01:50:54.242508Z digest=sha256:f3d76fe6b577b51ee3f2999ba2ed1f95165e5d5e36c5eab38e32a28246642049