Pith. sign in

Paper Citation Record · LEDGER

ChatSpot: Bootstrapping Multimodal LLMs via Precise Referring Instruction Tuning

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 9 inbound Pith citation observations for arXiv:2307.09474.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2307.09474 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 9 of 9 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:50:46.924144Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T01:27:31.611369Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 0ceb5c1d-5394-448c-84fb-587b59c230e3 · inbound

VModA: An Effective Framework for Adaptive NSFW Image Moderation cites this paper.

VModA: An Effective Framework for Adaptive NSFW Image Moderation ChatSpot: Bootstrapping Multimodal LLMs via Precise Referring Instruction Tuning

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T12:50:46.924144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:50:46.924144Z digest=sha256:9fe740db65ef20b299a2cf747bc2c88d6ead0c909781da42f52a7fb3a3470dae

Observation ccdc906f-94a4-4d76-bf34-91859924d5ba · inbound

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos cites this paper.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos ChatSpot: Bootstrapping Multimodal LLMs via Precise Referring Instruction Tuning

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:14.565949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:14.565949Z digest=sha256:61f4f2cf79e07ff1a071e901232b47a4d45b0e32138fa35320b52d85b3e2bc30

Observation 0aca91fb-18a3-4a47-8a79-f67907718a88 · inbound

Towards Better Dental AI: A Multimodal Benchmark and Instruction Dataset for Panoramic X-ray Analysis cites this paper.

Towards Better Dental AI: A Multimodal Benchmark and Instruction Dataset for Panoramic X-ray Analysis ChatSpot: Bootstrapping Multimodal LLMs via Precise Referring Instruction Tuning

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-04T19:27:40.876277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:27:40.876277Z digest=sha256:005b30711424203cb329bece0dfc27e34ac0618839887b041c1650e8eb991d45

Observation 1974987a-4acc-4cbd-b365-6465916833fb · inbound

LMMs Meet Object-Centric Vision: Understanding, Segmentation, Editing and Generation cites this paper.

LMMs Meet Object-Centric Vision: Understanding, Segmentation, Editing and Generation ChatSpot: Bootstrapping Multimodal LLMs via Precise Referring Instruction Tuning

Reference 236

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:11:08.680350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T15:35:37.095627Z digest=sha256:ffb3c72f0f4a59d249fbc4a35f7b7ffc097f1d562e379e97556007e9b08b7e14

Observation fefc4236-2a4c-457d-8fe8-4462ad955d71 · inbound

Injecting Distributional Awareness into MLLMs via Reinforcement Learning for Deep Imbalanced Regression cites this paper.

Injecting Distributional Awareness into MLLMs via Reinforcement Learning for Deep Imbalanced Regression ChatSpot: Bootstrapping Multimodal LLMs via Precise Referring Instruction Tuning

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T16:51:09.405573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-09T14:36:29.666730Z digest=sha256:9d3cba61dbe568cc3f3ed5fa440f9003f0672fb0922041b7fa913afba75928f2

Observation 1feea4e1-84f5-416f-8500-0c94edcb1002 · inbound

Injecting Distributional Awareness into MLLMs via Reinforcement Learning for Deep Imbalanced Regression cites this paper.

Injecting Distributional Awareness into MLLMs via Reinforcement Learning for Deep Imbalanced Regression ChatSpot: Bootstrapping Multimodal LLMs via Precise Referring Instruction Tuning

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T05:51:26.071245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-12T04:52:09.685243Z digest=sha256:cbc40eb907fbc4eee6ee9d074634b46ffe217a6c66aec90f140e62e540836b7c

Observation 099a5a36-851a-4396-b3f1-a86fa5d1d183 · inbound

See What I Mean: Aligning Vision and Language Representations for Video Fine-grained Object Understanding cites this paper.

See What I Mean: Aligning Vision and Language Representations for Video Fine-grained Object Understanding ChatSpot: Bootstrapping Multimodal LLMs via Precise Referring Instruction Tuning

Reference 110

Resolution
verified exact
arxiv_id, observed 2026-05-20T12:13:16.339503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-20T12:10:54.874012Z digest=sha256:9c35f4c1d137aabeb3f7cd7b98df3a064736b043f563a2964770978e92b8eaab

Observation 34bdb68c-e411-4eb2-9443-8e9093745aa2 · inbound

Vision Language Model Helps Private Information De-Identification in Vision Data cites this paper.

Vision Language Model Helps Private Information De-Identification in Vision Data ChatSpot: Bootstrapping Multimodal LLMs via Precise Referring Instruction Tuning

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T01:27:31.612626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-27T16:29:33.294960Z digest=sha256:e2d25b8f7f819c9c843d7fee369b3421520ab29ddb0133e5037e9f761dd890a7

Observation b035cd29-7572-42dd-a970-597b6f9d4499 · inbound

MentalThink: Shaping Thoughts in Mental SVG World cites this paper.

MentalThink: Shaping Thoughts in Mental SVG World ChatSpot: Bootstrapping Multimodal LLMs via Precise Referring Instruction Tuning

Reference 136

Resolution
unresolved
no resolver link, observed 2026-07-12T01:50:59.184754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-12T01:50:59.184754Z digest=sha256:3c025fdbad6d4931a5714fdc41a2c3a7000581e4e560f8421cba65f06d1307e5