Pith. sign in

Paper Citation Record · LEDGER

A Systematic Survey of Prompt Engineering on Vision-Language Foundation Models

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 17 inbound Pith citation observations for arXiv:2307.12980.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2307.12980 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 17 of 17 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 17 of 17 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:15:51.198362Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

63
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation f8af49c9-438f-4f61-8df0-50cd2d34fbae · inbound

Vibe Coding vs. Agentic Coding: Fundamentals and Practical Implications of Agentic AI cites this paper.

Vibe Coding vs. Agentic Coding: Fundamentals and Practical Implications of Agentic AI A Systematic Survey of Prompt Engineering on Vision-Language Foundation Models

Reference 203

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:51.198362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:51.198362Z digest=sha256:0014d334549429d34a8d25987a8407dfaf94f7b82f1c226c728bd35f7053db72

Observation 2fd5c4a3-f0f8-4472-b0af-28b870ad8721 · inbound

A Survey of Personalized Federated Foundation Models for Privacy-Preserving Recommendation cites this paper.

A Survey of Personalized Federated Foundation Models for Privacy-Preserving Recommendation A Systematic Survey of Prompt Engineering on Vision-Language Foundation Models

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-19T09:32:16.526928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-19T09:28:32.185398Z digest=sha256:7d98c61a2ba9e0b8f05516e75ada5e4c54072d40680b925e42388bfa8d725a29

Observation 482ae711-fa47-4911-ab40-5dd7c88d34a5 · inbound

Visual Textualization for Image Prompted Object Detection cites this paper.

Visual Textualization for Image Prompted Object Detection A Systematic Survey of Prompt Engineering on Vision-Language Foundation Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T21:37:04.064646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:37:04.064646Z digest=sha256:16b07c9609db4752f8b772319d32760bb05d600040e6ee1ee433cb66fcf0a4e1

Observation 26032027-4fc6-4c4b-95fc-8895bc238252 · inbound

A Survey of AIOps in the Era of Large Language Models cites this paper.

A Survey of AIOps in the Era of Large Language Models A Systematic Survey of Prompt Engineering on Vision-Language Foundation Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T23:26:36.620888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:26:36.620888Z digest=sha256:bc555f7f2151d2396a2ecc1d18f2700dd477b78a71d8c253f3fe48ea6818d5e7

Observation fe50a836-005b-4f99-9b63-5f8df3e74e80 · inbound

DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding cites this paper.

DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding A Systematic Survey of Prompt Engineering on Vision-Language Foundation Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T15:34:44.391857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:34:44.391857Z digest=sha256:d90894d4036a6abaf50309073a71ac7426ac22a38981c094fedddec54ae3c736

Observation 56635091-ef64-4704-98cf-59e530b62c46 · inbound

True Multimodal In-Context Learning Needs Attention to the Visual Context cites this paper.

True Multimodal In-Context Learning Needs Attention to the Visual Context A Systematic Survey of Prompt Engineering on Vision-Language Foundation Models

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-06T15:29:34.010676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:29:34.010676Z digest=sha256:26a95dbb78639fedf48818ddafae4ec1692f79d4674572c206ba5219f986faac

Observation a0d3e291-6774-4c4b-b33c-5f3ddc703e08 · inbound

Adapting Vision-Language Models Without Labels: A Comprehensive Survey cites this paper.

Adapting Vision-Language Models Without Labels: A Comprehensive Survey A Systematic Survey of Prompt Engineering on Vision-Language Foundation Models

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-05T23:17:06.585528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:17:06.585528Z digest=sha256:7de70f4c193c49dadcef31b42f130a242f0e57aba02771e2360014ec63c6e226

Observation 29bec3bd-5ffb-488e-a6ab-8b0bec84386a · inbound

CAD2DMD-SET: Synthetic Generation Tool of Digital Measurement Device CAD Model Datasets for fine-tuning Large Vision-Language Models cites this paper.

CAD2DMD-SET: Synthetic Generation Tool of Digital Measurement Device CAD Model Datasets for fine-tuning Large Vision-Language Models A Systematic Survey of Prompt Engineering on Vision-Language Foundation Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-05T14:06:18.622208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:06:18.622208Z digest=sha256:9fa0f0478462af6bfbb6ba6f296f571663e8b617af2db64b33db637865c8110b

Observation dbdf9168-c7a5-47e5-81e3-ca9122ddf634 · inbound

Modality Alignment across Trees on Heterogeneous Hyperbolic Manifolds cites this paper.

Modality Alignment across Trees on Heterogeneous Hyperbolic Manifolds A Systematic Survey of Prompt Engineering on Vision-Language Foundation Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T07:22:03.719205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:22:03.719205Z digest=sha256:bcbe11b361ae8eed4f3d3f258e896bd3e2590c824a06442f79307b49c8d8b52e

Observation 66fad044-bbab-42a8-bb45-6fae5efdc267 · inbound

AERMANI-VLM: Structured Prompting and Reasoning for Aerial Manipulation with Vision Language Models cites this paper.

AERMANI-VLM: Structured Prompting and Reasoning for Aerial Manipulation with Vision Language Models A Systematic Survey of Prompt Engineering on Vision-Language Foundation Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T00:24:38.922212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:24:38.922212Z digest=sha256:6ebc202cf2b3df7c1991aa010a73f2b6dec735eaf8e7f1b860e9426a065dec4a

Observation 76c436a8-ec59-4ec4-9f7c-4f762a8b5f5f · inbound

Are vision-language models ready to zero-shot replace supervised classification models in agriculture? cites this paper.

Are vision-language models ready to zero-shot replace supervised classification models in agriculture? A Systematic Survey of Prompt Engineering on Vision-Language Foundation Models

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-16T21:18:32.283981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-16T21:15:20.717705Z digest=sha256:5b5dd97326ddcc1d4f0e1d256b521e04526ba1b482fe736978b23c93c70e48cc

Observation 8fafcd6a-d497-4509-98ef-931aa2a2a01b · inbound

Does Your VFM Speak Plant? The Botanical Grammar of Vision Foundation Models for Object Detection cites this paper.

Does Your VFM Speak Plant? The Botanical Grammar of Vision Foundation Models for Object Detection A Systematic Survey of Prompt Engineering on Vision-Language Foundation Models

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-11T07:35:57.219371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T17:07:17.174815Z digest=sha256:ae066b300547a82f765224b47a13dbbc5134acba19e212c36d1177fcbd86a5e2

Observation cf9950a2-3e3e-40d7-b85a-7b2b6288e9c7 · inbound

Multilingual and Multimodal LLMs in the Wild: Building for Low-Resource Languages cites this paper.

Multilingual and Multimodal LLMs in the Wild: Building for Low-Resource Languages A Systematic Survey of Prompt Engineering on Vision-Language Foundation Models

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-20T14:38:21.886516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-20T14:33:36.100966Z digest=sha256:428cac1365cc3b98cf2d0c5a36a3349580ba87ea3c04c5199e9ad98c1015cb19

Observation 0f3cf654-188c-4da7-8cb4-4d8c14fd32b8 · inbound

PluRule: A Benchmark for Moderating Pluralistic Communities on Social Media cites this paper.

PluRule: A Benchmark for Moderating Pluralistic Communities on Social Media A Systematic Survey of Prompt Engineering on Vision-Language Foundation Models

Reference 176

Resolution
verified exact
arxiv_id, observed 2026-05-20T14:08:20.409821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-20T14:05:14.737146Z digest=sha256:1ec7650c401ecfcd2f530eb9506a5e4e539a511e2f71a64fb06aae8599ca09a9

Observation da5629af-2778-4f7e-9875-4d730b62fbeb · inbound

EPIG: Emotion-Based Prompting for Personalised Image Generation cites this paper.

EPIG: Emotion-Based Prompting for Personalised Image Generation A Systematic Survey of Prompt Engineering on Vision-Language Foundation Models

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-07-03T14:48:32.999944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T06:54:05.515501Z digest=sha256:0d0f5a14660dd6c49f0b48912aef655bc2c4b8c5f615e76ef8c19d0068b4759b

Observation 29510b11-2cee-45ee-927b-eff60a3f63ee · inbound

ZEBRA: Zero-Shot Entropy-Regularized Prompt Learning for Base-to-Novel Generalization in Audio-Language Models cites this paper.

ZEBRA: Zero-Shot Entropy-Regularized Prompt Learning for Base-to-Novel Generalization in Audio-Language Models A Systematic Survey of Prompt Engineering on Vision-Language Foundation Models

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T12:05:42.782901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-01T03:02:12.747334Z digest=sha256:85abccbb89d752b3c5cbc1d0cdaf55fbabd7ebb1af46d10f6846681314a0fc29

Observation b7cda6a1-eada-4c41-8416-e152951f0bc6 · inbound

Visual prompt engineering for video models cites this paper.

Visual prompt engineering for video models A Systematic Survey of Prompt Engineering on Vision-Language Foundation Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-01T02:13:12.198478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:13:12.198478Z digest=sha256:81be4a430375369f487a61c4ca823b1b029e59a6823dd3ca586ce6ad6d52ebf1