Pith. sign in

Paper Citation Record · LEDGER

A Survey of Vision-Language Pre-Trained Models

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 19 inbound Pith citation observations for arXiv:2202.10936.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2202.10936 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 19 of 19 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 19 of 19 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T17:57:33.435230Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T10:07:55.952015Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 17546764-88cc-4260-9f33-84695de99830 · inbound

Divide-and-Conquer: Tree-structured Strategy with Answer Distribution Estimator for Goal-Oriented Visual Dialogue cites this paper.

Divide-and-Conquer: Tree-structured Strategy with Answer Distribution Estimator for Goal-Oriented Visual Dialogue A Survey of Vision-Language Pre-Trained Models

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-08T17:57:33.435230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:57:33.435230Z digest=sha256:caa12dbaeba46f301f805e020ebe65b5bd35c0cc128d7ecf7ed957102dabe0ad

Observation 9cd3c01f-6c9a-4de7-9a68-3105ac6d19a2 · inbound

Mitigating Object Hallucination via Robust Local Perception Search cites this paper.

Mitigating Object Hallucination via Robust Local Perception Search A Survey of Vision-Language Pre-Trained Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T05:53:58.692319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:53:58.692319Z digest=sha256:352784bbe1466a5ce4fb9e079a3ed72bfb907954f11fcb0a5610b07c4ae397cd

Observation 13e46ba8-bdfe-4e22-aea4-0dc4a23c42a9 · inbound

MrM: Black-Box Membership Inference Attacks against Multimodal RAG Systems cites this paper.

MrM: Black-Box Membership Inference Attacks against Multimodal RAG Systems A Survey of Vision-Language Pre-Trained Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T05:41:16.296555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:41:16.296555Z digest=sha256:7cc9b96d3ffe34b1e902fa65a0beccc3826b45825e83dfea203d59cb0aee9ec0

Observation dfbd2932-fb48-49b2-b13d-92c274c9d958 · inbound

CF-VLM:CounterFactual Vision-Language Fine-tuning cites this paper.

CF-VLM:CounterFactual Vision-Language Fine-tuning A Survey of Vision-Language Pre-Trained Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:06.949603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:06.949603Z digest=sha256:6c9d6fd76170cc830ca6bbf7741fe1fe5f96aeacc040cb459005605a56acb3f5

Observation e4a9cd1b-e7b5-4309-9357-4c2ee836d135 · inbound

Generalizing vision-language models to novel domains: A comprehensive survey cites this paper.

Generalizing vision-language models to novel domains: A comprehensive survey A Survey of Vision-Language Pre-Trained Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:40.995438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:20:40.995438Z digest=sha256:4a79a5d5e6f99613573c3a8bc034e0747fca9180852fd88474c1d3314e5de06e

Observation ffdce5f9-038f-4fa9-83e1-578385bda8de · inbound

LEGO Co-builder: Exploring Fine-Grained Vision-Language Modeling for Multimodal LEGO Assembly Assistants cites this paper.

LEGO Co-builder: Exploring Fine-Grained Vision-Language Modeling for Multimodal LEGO Assembly Assistants A Survey of Vision-Language Pre-Trained Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T19:29:39.359722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:29:39.359722Z digest=sha256:5718e83c036ee9a13a6e81546b43e4ce96d4775730057e083386a4012b1f9751

Observation 186cde67-7b07-4077-919c-ee7175e5e096 · inbound

Innocence in the Crossfire: Roles of Skip Connections in Jailbreaking Visual Language Models cites this paper.

Innocence in the Crossfire: Roles of Skip Connections in Jailbreaking Visual Language Models A Survey of Vision-Language Pre-Trained Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T16:25:54.920538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:25:54.920538Z digest=sha256:4c1fef23a0443136be6ffb6864489624461f27d4e1344d7b13b31015ff9b5f1d

Observation 902feb0e-ce7e-4d32-bab5-23c0b610e022 · inbound

Continual Learning for VLMs: A Survey and Taxonomy Beyond Forgetting cites this paper.

Continual Learning for VLMs: A Survey and Taxonomy Beyond Forgetting A Survey of Vision-Language Pre-Trained Models

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-21T23:44:26.880553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-21T23:41:54.385864Z digest=sha256:f97c9ec558fb182740218096184a7d5dbfc8f75b3c0ab95e3babb749c89e9bfb

Observation 36838e3d-44fc-4ce7-956e-97a6ccccf560 · inbound

Continual Learning for VLMs: A Survey and Taxonomy Beyond Forgetting cites this paper.

Continual Learning for VLMs: A Survey and Taxonomy Beyond Forgetting A Survey of Vision-Language Pre-Trained Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T00:51:52.046324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:51:52.046324Z digest=sha256:967cc59d2b6ab60ae66f24ec7d89605a2d5427f1b51f67cff1c9a551b2e6bd4b

Observation 3e761e56-ebad-4d17-ac87-ec51b196b33a · inbound

Training-Free Multimodal Large Language Model Orchestration cites this paper.

Training-Free Multimodal Large Language Model Orchestration A Survey of Vision-Language Pre-Trained Models

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-19T00:12:54.098165Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-19T00:12:39.834892Z digest=sha256:ae3656b8b9b04c44528ad19f13a429b634f823d0eb22b9013101071b3a6c00f2

Observation 65e4103a-0287-481d-8498-e29585f1edac · inbound

Training-Free Multimodal Large Language Model Orchestration cites this paper.

Training-Free Multimodal Large Language Model Orchestration A Survey of Vision-Language Pre-Trained Models

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-25T08:05:30.720002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-25T08:02:15.950975Z digest=sha256:4dcf2b6a36cefa7622c8fd9dff41382472ece1f5093add3e3ae8c7c491aa4e0e

Observation f339b4fc-c974-4f55-a6c2-54a2ebadcc84 · inbound

Temporal Inversion for Learning Interval Change in Chest X-Rays cites this paper.

Temporal Inversion for Learning Interval Change in Chest X-Rays A Survey of Vision-Language Pre-Trained Models

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:00:50.514589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T19:23:46.184189Z digest=sha256:dc2da05a0bf1a068edf18ab6349035a99995785b998df1a02cc2c38290f5e4c6

Observation 24eb3a63-25db-4324-aa8a-eb97406f0233 · inbound

Towards Long-horizon Agentic Multimodal Search cites this paper.

Towards Long-horizon Agentic Multimodal Search A Survey of Vision-Language Pre-Trained Models

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:01:04.591634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T15:40:32.137708Z digest=sha256:d2f29d8c8084318fcda2d996fc16ae2d5fbaa509fed3b3454352273ae2b5ad7c

Observation 3bc4273e-a612-4460-9f37-0f2a9c09dec1 · inbound

Federated Cross-Modal Retrieval with Missing Modalities via Semantic Routing and Adapter Personalization cites this paper.

Federated Cross-Modal Retrieval with Missing Modalities via Semantic Routing and Adapter Personalization A Survey of Vision-Language Pre-Trained Models

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:06:09.279154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-08T12:44:55.026879Z digest=sha256:2e094ec44189b5906b3c700ef8f11727617736715c2445a13bb807699c7039af

Observation a666579a-421f-4350-b4ae-ae2ab06de875 · inbound

Structural Ranking of the Cognitive Plausibility of Computational Models of Analogy and Metaphors with the Minimal Cognitive Grid cites this paper.

Structural Ranking of the Cognitive Plausibility of Computational Models of Analogy and Metaphors with the Minimal Cognitive Grid A Survey of Vision-Language Pre-Trained Models

Reference 50

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T16:56:06.216704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-09T14:33:11.033906Z digest=sha256:f422611426196804d8e6a7863398e0228bb4ce90f4216ac007922c941710e781

Observation cd80b9d0-947a-4308-b686-af6fbeebf19a · inbound

Efficient Prompt Learning for Traffic Forecasting cites this paper.

Efficient Prompt Learning for Traffic Forecasting A Survey of Vision-Language Pre-Trained Models

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T08:06:26.722342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-12T01:20:37.192907Z digest=sha256:56eb44c8b1067384b1ac3d9c2b205c4ed7379aef2b8ade5fc041b329b241087d

Observation 1ef376ca-e6ba-4686-8f92-8eba6c2e91fd · inbound

Horizontal and Longitudinal Comparisons Among AI Subfields: A Bibliometric Perspective cites this paper.

Horizontal and Longitudinal Comparisons Among AI Subfields: A Bibliometric Perspective A Survey of Vision-Language Pre-Trained Models

Reference 35

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T07:56:28.056860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-12T01:29:18.418975Z digest=sha256:05fd4883fdf64839dca274374f7c67b033fa42bf6946281e941d11fd1df5762a

Observation 10f370ce-81f6-4164-bfa4-29f6e9f9f733 · inbound

TouchThinker: Scaling Tactile Commonsense Reasoning to the Open World with Large-scale Data and Action-aware Representation cites this paper.

TouchThinker: Scaling Tactile Commonsense Reasoning to the Open World with Large-scale Data and Action-aware Representation A Survey of Vision-Language Pre-Trained Models

Reference 25

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T10:07:55.953745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-27T10:15:59.516498Z digest=sha256:d53ff8ab571b6d39970efb64a8b9c003e072a0f7b2c2b3360ff3177ac7add880

Observation cecabeb0-c43d-4b02-903b-316868b99a22 · inbound

TouchThinker: Scaling Tactile Commonsense Reasoning to the Open World with Large-scale Data and Action-aware Representation cites this paper.

TouchThinker: Scaling Tactile Commonsense Reasoning to the Open World with Large-scale Data and Action-aware Representation A Survey of Vision-Language Pre-Trained Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-14T18:06:38.115031Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T18:06:38.115031Z digest=sha256:6dd52143c30739f54f51938bfb71a1b208c8ce6092abe9dbbe53acb1ed831b61