Pith. sign in

Paper Citation Record · LEDGER

VISA: Reasoning Video Object Segmentation via Large Language Models

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 6 inbound Pith citation observations for arXiv:2407.11325.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2407.11325 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 6 of 6 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:02:52.905262Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-23T04:47:33.548752Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 9938afde-aaa0-4c5f-9cfd-741d2486861c · inbound

Efficient Reasoning with Hidden Thinking cites this paper.

Efficient Reasoning with Hidden Thinking VISA: Reasoning Video Object Segmentation via Large Language Models

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-23T04:47:33.552448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-23T04:45:38.608009Z digest=sha256:dc9b3ad4a8dd96a4e4f6afdce8d834228347c4efd0c3c2fac456bc93c32cd93b

Observation 7dc7196d-1d7d-4a5b-9b90-e1fa8f024dc9 · inbound

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration cites this paper.

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration VISA: Reasoning Video Object Segmentation via Large Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:52.905262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:52.905262Z digest=sha256:788ffd79294a60b85edaf626c97443c8104b56f4b4d4c96e32268b665ba1451b

Observation 9ea3d715-7a0c-43e1-96a9-31bf0a037463 · inbound

Unleashing Hierarchical Reasoning: An LLM-Driven Framework for Training-Free Referring Video Object Segmentation cites this paper.

Unleashing Hierarchical Reasoning: An LLM-Driven Framework for Training-Free Referring Video Object Segmentation VISA: Reasoning Video Object Segmentation via Large Language Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-05T05:09:31.758998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T05:09:31.758998Z digest=sha256:0fcddd6e1b4f04f5d4019d414110b0e0755cfd0b1d390960c425d4f725adb647

Observation ccd7c3eb-cf59-4a92-b883-3e796c194a95 · inbound

APRVOS: 1st Place Winner of 5th PVUW MeViS-Audio Track cites this paper.

APRVOS: 1st Place Winner of 5th PVUW MeViS-Audio Track VISA: Reasoning Video Object Segmentation via Large Language Models

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-10T03:29:21.548072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T03:29:12.783993Z digest=sha256:b73cf51e033ceb324b8cc9b4616fe4df39364547d2ad6de2c0cbf6437d5b3ce5

Observation eff002f5-1923-4ad0-853c-d032ceecd1ff · inbound

Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model cites this paper.

Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model VISA: Reasoning Video Object Segmentation via Large Language Models

Reference 159

Resolution
unresolved
no resolver link, observed 2026-07-31T06:20:14.277951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:20:14.277951Z digest=sha256:981e842d52704b7f87ef044c6a6c3a6abcbc187fb5dbe9540b0d5ec6e4c77d2d

Observation 326989ef-ee08-4941-a8c3-e4ee0f13e581 · inbound

Beyond Frame Selection: Generative Latent Evidence Aggregation for Long-Video Understanding cites this paper.

Beyond Frame Selection: Generative Latent Evidence Aggregation for Long-Video Understanding VISA: Reasoning Video Object Segmentation via Large Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-07-31T04:55:07.042876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T04:55:07.042876Z digest=sha256:ecaa4faa225673aea006e0081cf36875faa84fe75f817ff74b065d624fbc353b