Pith. sign in

Paper Citation Record · LEDGER

LLaVA-Interactive: An All-in-One Demo for Image Chat, Segmentation, Generation and Editing

As of 17 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 15 inbound Pith citation observations for arXiv:2311.00571.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2311.00571 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 15 of 15 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 15 of 15 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T05:46:49.011495Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-30T19:15:01.025528Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation eb25dc0a-a158-4f21-ae37-4cfd71f0dcb7 · inbound

Document Parsing Unveiled: Techniques, Challenges, and Prospects for Structured Information Extraction cites this paper.

Document Parsing Unveiled: Techniques, Challenges, and Prospects for Structured Information Extraction LLaVA-Interactive: An All-in-One Demo for Image Chat, Segmentation, Generation and Editing

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-23T19:15:47.116632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-23T19:15:21.695801Z digest=sha256:1bdfe57e0ff0b593b0d6505eadf889e4b4cc1fcb2b6493a279b7cd2bd0c31272

Observation ba40203d-e31d-4495-a19d-ccfea8043e4f · inbound

SPAgent: Adaptive Task Decomposition and Model Selection for General Video Generation and Editing cites this paper.

SPAgent: Adaptive Task Decomposition and Model Selection for General Video Generation and Editing LLaVA-Interactive: An All-in-One Demo for Image Chat, Segmentation, Generation and Editing

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T10:44:14.694941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:44:14.694941Z digest=sha256:e91920fa9cbfaf5ab271b077c8aead6f05a9f8897380ac1456d48005ff0f73c3

Observation 66c596f8-bef6-47a7-92f4-f7f6df5ca87c · inbound

Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos cites this paper.

Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos LLaVA-Interactive: An All-in-One Demo for Image Chat, Segmentation, Generation and Editing

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-16T11:39:22.536662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-16T11:39:22.340737Z digest=sha256:8d0615922ebd110e4a8d1158e0a13498be56ce710543549733b6ffa28ffc7127

Observation 1b14393b-8007-46ad-8a5f-cb8277c9670f · inbound

Advancing General Multimodal Capability of Vision-language Models with Pyramid-descent Visual Position Encoding cites this paper.

Advancing General Multimodal Capability of Vision-language Models with Pyramid-descent Visual Position Encoding LLaVA-Interactive: An All-in-One Demo for Image Chat, Segmentation, Generation and Editing

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T18:53:21.142546Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T18:53:21.142546Z digest=sha256:a59243036431d873785417ffdd3209b9bb21d958a9b823647f3489a8aacabbd9

Observation 7f851866-9bdf-43f3-a7f7-00b6dc322be3 · inbound

Foundation Model-Driven Framework for Human-Object Interaction Prediction with Segmentation Mask Integration cites this paper.

Foundation Model-Driven Framework for Human-Object Interaction Prediction with Segmentation Mask Integration LLaVA-Interactive: An All-in-One Demo for Image Chat, Segmentation, Generation and Editing

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-16T05:46:49.011495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:46:49.011495Z digest=sha256:eafe5b81e215d76f21a5ae1baa019cb3fbc34ca5cc9b5815d83f4247f828ae9a

Observation 5297fe55-ba3d-404c-868d-b4dcd4b96b06 · inbound

MCCD: Multi-Agent Collaboration-based Compositional Diffusion for Complex Text-to-Image Generation cites this paper.

MCCD: Multi-Agent Collaboration-based Compositional Diffusion for Complex Text-to-Image Generation LLaVA-Interactive: An All-in-One Demo for Image Chat, Segmentation, Generation and Editing

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-16T00:49:40.189040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:49:40.189040Z digest=sha256:60ee926dbe2a8391b076be8c29d0b63cc9d3ddfd3ed031c3b4b104c343fbd44f

Observation f4d3b2e0-f4c1-4ec3-9c84-aefc544a0028 · inbound

ZeroVO: Visual Odometry with Minimal Assumptions cites this paper.

ZeroVO: Visual Odometry with Minimal Assumptions LLaVA-Interactive: An All-in-One Demo for Image Chat, Segmentation, Generation and Editing

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T05:26:58.939069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:26:58.939069Z digest=sha256:318646ec2eb54887b70a6caf98b13ca275389ff73e2f659da77a5ce5e9ba294d

Observation 3110745d-1193-4f13-a9bd-f888d9d002cd · inbound

Training-Free Multimodal Large Language Model Orchestration cites this paper.

Training-Free Multimodal Large Language Model Orchestration LLaVA-Interactive: An All-in-One Demo for Image Chat, Segmentation, Generation and Editing

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-19T00:12:54.027421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-19T00:12:39.834892Z digest=sha256:6097ec2d0e74ce9a01ee3f555624c2555bff849a64bd77e14f980fa3d6d53d26

Observation 49da9031-fc32-4f2b-a3b2-c10b26e3ddb5 · inbound

Training-Free Multimodal Large Language Model Orchestration cites this paper.

Training-Free Multimodal Large Language Model Orchestration LLaVA-Interactive: An All-in-One Demo for Image Chat, Segmentation, Generation and Editing

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-25T08:05:30.658939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-25T08:02:15.950975Z digest=sha256:996c8ce0190d5c75fe976c4e61a8024ec27186c96d303ccb4c779221ef75a1fd

Observation 0fbf6a59-969d-4688-8b4e-5a562fb3d538 · inbound

Empowering Multimodal LLMs with External Tools: A Comprehensive Survey cites this paper.

Empowering Multimodal LLMs with External Tools: A Comprehensive Survey LLaVA-Interactive: An All-in-One Demo for Image Chat, Segmentation, Generation and Editing

Reference 273

Resolution
unresolved
no resolver link, observed 2026-08-05T20:29:09.987685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:29:09.987685Z digest=sha256:695f8ecb39b4cb632db43e9a76fad0f1f4f9ec70283497e6d4c96fa663dd918f

Observation 79d84097-cb4d-4ab8-b0a8-74a8551b01ba · inbound

Ego-centric Predictive Model Conditioned on Hand Trajectories cites this paper.

Ego-centric Predictive Model Conditioned on Hand Trajectories LLaVA-Interactive: An All-in-One Demo for Image Chat, Segmentation, Generation and Editing

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T15:29:25.367304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:29:25.367304Z digest=sha256:98fddb4f0913ec7f5c1558eb2a8e37a0a5a4cf2dece90f110ae3f3caaca3458f

Observation 90508fbc-5b49-4985-b156-1603ded0066c · inbound

Relightable Gaussian Splatting for Virtual Production Using Image-Based Illumination cites this paper.

Relightable Gaussian Splatting for Virtual Production Using Image-Based Illumination LLaVA-Interactive: An All-in-One Demo for Image Chat, Segmentation, Generation and Editing

Reference 49

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T07:46:35.198556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-05-12T01:56:34.807815Z digest=sha256:ea473d6cacabc6a92b304b0c71e67008b74a10f9c03b36e68253e9c4b0aa6eb8

Observation fefe5fc2-e163-4c81-9296-2c7661de7aef · inbound

Don't Guess, Just Ask: Resolving Ambiguity in Referring Segmentation via Multi-turn Clarification cites this paper.

Don't Guess, Just Ask: Resolving Ambiguity in Referring Segmentation via Multi-turn Clarification LLaVA-Interactive: An All-in-One Demo for Image Chat, Segmentation, Generation and Editing

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-20T13:43:19.800281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T13:38:21.492236Z digest=sha256:8a5c2dda928ae90fae24db6c21d9d360b999dc2fd7bd4509726708d03d931622

Observation c7e300d8-00ce-458c-96fc-2b98f9977063 · inbound

Don't Guess, Just Ask: Resolving Ambiguity in Referring Segmentation via Multi-turn Clarification cites this paper.

Don't Guess, Just Ask: Resolving Ambiguity in Referring Segmentation via Multi-turn Clarification LLaVA-Interactive: An All-in-One Demo for Image Chat, Segmentation, Generation and Editing

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-06-30T19:15:01.027802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-30T19:07:05.208780Z digest=sha256:9f83dce84bf55506974c80b6f82dd4ce9a334a3b5e3f73e89cfe0918e89af5f5

Observation 9d9f1967-9619-4e5d-b923-c34257ae8a5a · inbound

A Comprehensive Study of Implementation Bugs in Multi-modal Agents cites this paper.

A Comprehensive Study of Implementation Bugs in Multi-modal Agents LLaVA-Interactive: An All-in-One Demo for Image Chat, Segmentation, Generation and Editing

Reference 54

Resolution
unresolved
no resolver link, observed 2026-07-11T10:41:33.954036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T10:41:33.954036Z digest=sha256:362d7b4f71a3cc97f524feb3961fce63340b6ba3be5b23636a7040185dc2df1f