Pith. sign in

Paper Citation Record · LEDGER

LLaVA-Plus: Learning to Use Tools for Creating Multimodal Agents

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 23 inbound Pith citation observations for arXiv:2311.05437.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2311.05437 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 23 of 23 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 23 of 23 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T19:25:33.751527Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T10:35:41.926606Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation b96c18cb-aa7d-4802-81af-1c9b31884c8a · inbound

MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices cites this paper.

MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices LLaVA-Plus: Learning to Use Tools for Creating Multimodal Agents

Reference 78

Resolution
verified exact
arxiv_id, observed 2026-05-16T16:35:38.185520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T16:35:37.937462Z digest=sha256:1e20ceb84a222da8220b4f281a1f98635360cd6ac4c460adf869c7f7bb5d4a7f

Observation 4401e72f-91e8-4b34-a444-fad6f2db5712 · inbound

A Survey on Hallucination in Large Vision-Language Models cites this paper.

A Survey on Hallucination in Large Vision-Language Models LLaVA-Plus: Learning to Use Tools for Creating Multimodal Agents

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-13T22:10:10.338212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-13T22:10:10.186950Z digest=sha256:b31b6105df49b60230ef41eaa8f65f519ad9c52e5589117d187dd9c8be5c7861

Observation ffb75772-aa49-4876-8637-f0066da0672e · inbound

Large Language Models: A Survey cites this paper.

Large Language Models: A Survey LLaVA-Plus: Learning to Use Tools for Creating Multimodal Agents

Reference 217

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:22:55.992348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T15:22:54.023279Z digest=sha256:1c56641ff639717dd6c7c5e86580260b081ae305267008bc2f1c2b275bd2a2c6

Observation 1ffe80b4-f047-46b4-9bbf-997312599f30 · inbound

MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training cites this paper.

MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training LLaVA-Plus: Learning to Use Tools for Creating Multimodal Agents

Reference 77

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T04:09:36.536840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T04:09:36.019146Z digest=sha256:3afcf1943b56103a2fb2802cea5bfb063bc6f1c4b68f2104413d835eab9e653a

Observation 66ac61e0-3ecd-462d-97f2-86c17d6da27b · inbound

Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model cites this paper.

Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model LLaVA-Plus: Learning to Use Tools for Creating Multimodal Agents

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:57:26.964174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-13T05:57:26.887069Z digest=sha256:e91ebe6f569ca6b56c0feec42502d3dc4aaa3454943f5f63b34fdc15d7b22c4b

Observation fd30c3a7-9fb9-4fe2-b9b9-cc64a6e87e46 · inbound

Document Parsing Unveiled: Techniques, Challenges, and Prospects for Structured Information Extraction cites this paper.

Document Parsing Unveiled: Techniques, Challenges, and Prospects for Structured Information Extraction LLaVA-Plus: Learning to Use Tools for Creating Multimodal Agents

Reference 136

Resolution
verified exact
arxiv_id, observed 2026-05-23T19:15:47.062053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-23T19:15:21.695801Z digest=sha256:f9143f4ec9bbb7d21c4a3e1d4381d034f0ab447f32dd11291fc5f84de68f51ea

Observation c09a4593-00e2-4dac-b352-4eecbd225704 · inbound

Image Embedding Sampling Method for Diverse Captioning cites this paper.

Image Embedding Sampling Method for Diverse Captioning LLaVA-Plus: Learning to Use Tools for Creating Multimodal Agents

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T19:25:33.751527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:25:33.751527Z digest=sha256:88e219e6939e233dce7f16e1ef7df8e3f98270f6f8b3bd20f947e8d5dedf019e

Observation 42de5efa-0cb5-431a-83ba-d78fd3927afc · inbound

Circle-RoPE: Cone-like Decoupled Rotary Positional Embedding for Large Vision-Language Models cites this paper.

Circle-RoPE: Cone-like Decoupled Rotary Positional Embedding for Large Vision-Language Models LLaVA-Plus: Learning to Use Tools for Creating Multimodal Agents

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-22T14:21:39.732552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T14:19:34.622854Z digest=sha256:73aa6f0fc4d272349a5d780f79e7259d239761f4d74a7b7e0425edd78251dd0e

Observation 5d56ce10-92a8-4789-9fc0-94d03b860da1 · inbound

Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought cites this paper.

Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought LLaVA-Plus: Learning to Use Tools for Creating Multimodal Agents

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T12:42:23.566141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:42:23.566141Z digest=sha256:774213324786208773591f1e0069b71724a4ef311b762debb251ef0cd2bf7640

Observation b0e4dc90-d51d-4690-9e97-909398100931 · inbound

ZeroVO: Visual Odometry with Minimal Assumptions cites this paper.

ZeroVO: Visual Odometry with Minimal Assumptions LLaVA-Plus: Learning to Use Tools for Creating Multimodal Agents

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T05:26:59.109392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:26:59.109392Z digest=sha256:e7379285474748c716e7f495bddbeb076283d81658811be57d8d2110b8cf12c6

Observation 4720f529-b09c-480a-a074-accdfeef8031 · inbound

A Comprehensive Survey of Deep Research: Systems, Methodologies, and Applications cites this paper.

A Comprehensive Survey of Deep Research: Systems, Methodologies, and Applications LLaVA-Plus: Learning to Use Tools for Creating Multimodal Agents

Reference 152

Resolution
unresolved
no resolver link, observed 2026-08-07T00:48:20.777371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:48:20.777371Z digest=sha256:d7f62d485cd165580f74f3df8bd347022406b35d87d3fece9f7e6a01f020a006

Observation b6798a7d-cf83-4917-bc88-4894cf994b78 · inbound

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance cites this paper.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance LLaVA-Plus: Learning to Use Tools for Creating Multimodal Agents

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T19:25:02.756176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:25:02.756176Z digest=sha256:bb8b98d4e1714a17cc2139dff819b36252ce39b5dcde29f8822156bd4e9b2847

Observation 13d2176b-db56-4ac4-856e-40753e4f798d · inbound

Augmented Vision-Language Models: A Systematic Review cites this paper.

Augmented Vision-Language Models: A Systematic Review LLaVA-Plus: Learning to Use Tools for Creating Multimodal Agents

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-06T14:33:37.251420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:33:37.251420Z digest=sha256:2c4a5504139d93595fc85e2c7bcb8891d2be7b48c540c230fc6ee46e76d7453d

Observation 3665d3b9-926c-4d2e-9dea-54a19b7b8306 · inbound

Reinforced Visual Perception with Tools cites this paper.

Reinforced Visual Perception with Tools LLaVA-Plus: Learning to Use Tools for Creating Multimodal Agents

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-05T12:27:04.989084Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T12:27:04.989084Z digest=sha256:a6b0bd9e6ef92bffa802820c8529b7d83e157c1bd4ec9dff56ef128d1c5db4ef

Observation ae82f0c8-2299-4db5-8847-af159577028b · inbound

Mitigating Coordinate Prediction Bias from Positional Encoding Failures cites this paper.

Mitigating Coordinate Prediction Bias from Positional Encoding Failures LLaVA-Plus: Learning to Use Tools for Creating Multimodal Agents

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T05:12:23.519700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-18T05:11:30.734633Z digest=sha256:9fa9d0ddd84cd9c5a41e66698b041b153f7cbc3e2b2307fd6697b7730b99491f

Observation df93be91-3232-4a8c-b30b-18c5ae5848ed · inbound

VisReason: A Large-Scale Dataset for Visual Chain-of-Thought Reasoning cites this paper.

VisReason: A Large-Scale Dataset for Visual Chain-of-Thought Reasoning LLaVA-Plus: Learning to Use Tools for Creating Multimodal Agents

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-03T20:57:30.428570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:57:30.428570Z digest=sha256:df0e576d38af263f72a44f895443e40d7db12a6b3d14345251c0c853c2a4a5af

Observation 8089de90-037a-42c6-8409-83dbeba083cf · inbound

E-VLA: Event-Augmented Vision-Language-Action Model for Dark and Blurred Scenes cites this paper.

E-VLA: Event-Augmented Vision-Language-Action Model for Dark and Blurred Scenes LLaVA-Plus: Learning to Use Tools for Creating Multimodal Agents

Reference 25

Resolution
unresolved
no resolver link, observed 2026-07-13T09:40:35.631188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T09:40:35.631188Z digest=sha256:f0960e6e441c154531510ab869bd6bac4402cdddc3bf443b0bae673641ff856d

Observation a3d56488-2b30-43f2-b2fe-443ab946a08e · inbound

Less Detail, Better Answers: Degradation-Driven Prompting for VQA cites this paper.

Less Detail, Better Answers: Degradation-Driven Prompting for VQA LLaVA-Plus: Learning to Use Tools for Creating Multimodal Agents

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:05:47.973219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T20:17:01.867903Z digest=sha256:c9cc34530517511db2f90556694c097a20170068c0c042b543118abcd1ec68d8

Observation ac93bcb1-9b2b-44dc-acb9-751745bcae77 · inbound

Waking Up Blind: Cold-Start Optimization of Supervision-Free Agentic Trajectories for Grounded Visual Perception cites this paper.

Waking Up Blind: Cold-Start Optimization of Supervision-Free Agentic Trajectories for Grounded Visual Perception LLaVA-Plus: Learning to Use Tools for Creating Multimodal Agents

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-10T05:36:01.721126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T05:33:33.589721Z digest=sha256:bf5c9f9b1d566d18f0066b17b0619f50ffb68be163052bfb61d11e84d3e0ea2b

Observation b24b2b45-f5a1-4dce-8f76-ad08bd6af5a0 · inbound

Second-Order Multi-Level Variance Correction for Modality Competition in Multimodal Models cites this paper.

Second-Order Multi-Level Variance Correction for Modality Competition in Multimodal Models LLaVA-Plus: Learning to Use Tools for Creating Multimodal Agents

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-20T19:03:39.571289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T19:02:02.266052Z digest=sha256:7d8a443e6c50d174a08aa0c162a2aee4b032a236b387cc9989abd8e51b5e69c4

Observation 45a2a72a-1e30-471f-b9a5-d41b606ec131 · inbound

IndusAgent: Reinforcing Open-Vocabulary Industrial Anomaly Detection with Agentic Tools cites this paper.

IndusAgent: Reinforcing Open-Vocabulary Industrial Anomaly Detection with Agentic Tools LLaVA-Plus: Learning to Use Tools for Creating Multimodal Agents

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-21T05:33:58.501835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T05:33:30.670201Z digest=sha256:90572256bd8c021d82132e6cbb093b8f23e7dcea3e05fde977cbb3b1ef77f185

Observation d590c960-c4d2-457b-942d-8adab7f652f6 · inbound

ReGRPO: Reflection-Augmented Policy Optimization for Tool-Using Agents cites this paper.

ReGRPO: Reflection-Augmented Policy Optimization for Tool-Using Agents LLaVA-Plus: Learning to Use Tools for Creating Multimodal Agents

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T10:35:41.928986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-01T05:23:21.162004Z digest=sha256:63a0d7b48b733ace9289a21b52838d906dd347a73cd44e208ff559b8968b8c1a

Observation 61b9a200-a680-4ee8-b4db-2cd88dac3504 · inbound

BVS: Bayesian Visual Search with Multimodal Large Language Model for Fine-grained Perception cites this paper.

BVS: Bayesian Visual Search with Multimodal Large Language Model for Fine-grained Perception LLaVA-Plus: Learning to Use Tools for Creating Multimodal Agents

Reference 216

Resolution
unresolved
no resolver link, observed 2026-07-12T04:17:40.198357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-12T04:17:40.198357Z digest=sha256:11048ef522a08ec89f3647b3a3898979cb49c48c88537b5be9f2f8bf3641fc58