Pith. sign in

Paper Citation Record · LEDGER

ChatGPT Asks, BLIP-2 Answers: Automatic Questioning Towards Enriched Visual Descriptions

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 8 inbound Pith citation observations for arXiv:2303.06594.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2303.06594 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 8 of 8 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T00:23:32.836660Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-17T14:22:18.666015Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation f85ff31d-71c3-43bd-99a0-5650941854ba · inbound

MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models cites this paper.

MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models ChatGPT Asks, BLIP-2 Answers: Automatic Questioning Towards Enriched Visual Descriptions

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-10T20:37:01.800800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T20:37:01.617345Z digest=sha256:353234a7ee2109ecc1aefc785dd0abde32bc147c470f5c9ee7d0a39c6bf6174b

Observation 4fac5e0e-1e55-4b6f-8834-4bcd53a90ae5 · inbound

A Survey on Multimodal Large Language Models cites this paper.

A Survey on Multimodal Large Language Models ChatGPT Asks, BLIP-2 Answers: Automatic Questioning Towards Enriched Visual Descriptions

Reference 82

Resolution
verified exact
arxiv_id, observed 2026-05-16T02:56:42.603376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T02:56:41.658658Z digest=sha256:1f95dbd42d5cc9a71b641209e0af81bd3a0e75fcd4c6533843fc6722d65eb735

Observation b0407eb2-6147-421f-ae70-a19b3137e8b1 · inbound

MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning cites this paper.

MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning ChatGPT Asks, BLIP-2 Answers: Automatic Questioning Towards Enriched Visual Descriptions

Reference 59

Resolution
verified exact
arxiv_id, observed 2026-05-16T07:13:09.013573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T07:13:08.867745Z digest=sha256:92f4bacf5a8cc66ba2fd01ae7fbdaa46dabe1190ed352b500f8397db3021541c

Observation 658192d9-7847-432d-8a7e-57eada58229f · inbound

An Embodied Generalist Agent in 3D World cites this paper.

An Embodied Generalist Agent in 3D World ChatGPT Asks, BLIP-2 Answers: Automatic Questioning Towards Enriched Visual Descriptions

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-17T14:22:18.668885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T14:22:18.606817Z digest=sha256:1dbf9d74450b3beff690aaed345f507c583eb1af1efa713606ca68d1bbc1f235

Observation 7d012bea-a758-4e5b-b7e6-288bc1cba0b5 · inbound

Adapting Lightweight Vision Language Models for Radiological Visual Question Answering cites this paper.

Adapting Lightweight Vision Language Models for Radiological Visual Question Answering ChatGPT Asks, BLIP-2 Answers: Automatic Questioning Towards Enriched Visual Descriptions

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T00:23:32.836660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:23:32.836660Z digest=sha256:b4482324a90f50bd07cc59a0506f6a7159f2b33348255c509940f8d69120fe70

Observation 5afc66ba-1db5-4322-8cac-109c8ec26260 · inbound

ReME: A Data-Centric Framework for Training-Free Open-Vocabulary Segmentation cites this paper.

ReME: A Data-Centric Framework for Training-Free Open-Vocabulary Segmentation ChatGPT Asks, BLIP-2 Answers: Automatic Questioning Towards Enriched Visual Descriptions

Reference 88

Resolution
malformed identifier
no resolver link, observed 2026-08-06T22:36:50.080336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:50.080336Z digest=sha256:cba46140b9c240da5362b09a9c8311a53487e008e21925d8088648f00d18b52e

Observation 693d6196-8f77-4f8b-8a6c-a0382fd50ecd · inbound

Prompt Engineering in Segment Anything Model: Methodologies, Applications, and Emerging Challenges cites this paper.

Prompt Engineering in Segment Anything Model: Methodologies, Applications, and Emerging Challenges ChatGPT Asks, BLIP-2 Answers: Automatic Questioning Towards Enriched Visual Descriptions

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-06T17:56:03.863141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:56:03.863141Z digest=sha256:19820807dfab1da864b324ac25f7623923a9b58b115f8c2add830a56f5d197f5

Observation ac2d11c2-308f-4e49-badc-dc4acc1c35df · inbound

MM-Nav: Multi-View VLA Model for Robust Visual Navigation via Multi-Expert Learning cites this paper.

MM-Nav: Multi-View VLA Model for Robust Visual Navigation via Multi-Expert Learning ChatGPT Asks, BLIP-2 Answers: Automatic Questioning Towards Enriched Visual Descriptions

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-04T12:37:55.687390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:37:55.687390Z digest=sha256:cddc7571e8b309b30d75ced34e60eb4f8f506e5b9241c1735094ec478165cf66