Pith. sign in

Paper Citation Record · LEDGER

Multi-Grained Vision Language Pre-Training: Aligning Texts with Visual Concepts

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 17 inbound Pith citation observations for arXiv:2111.08276.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2111.08276 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 17 of 17 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 17 of 17 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T05:30:25.558782Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T17:18:43.859907Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 81367072-436f-4468-a3ee-66063c46e2dc · inbound

ViperGPT: Visual Inference via Python Execution for Reasoning cites this paper.

ViperGPT: Visual Inference via Python Execution for Reasoning Multi-Grained Vision Language Pre-Training: Aligning Texts with Visual Concepts

Reference 65

Resolution
verified exact
arxiv_id, observed 2026-05-17T18:15:14.683921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T18:15:14.382011Z digest=sha256:72c20910db5b59a1654863971286fc625c842281c16d7b3068a4cc15579a1e53

Observation 1f50a754-badd-479f-98f7-5b14e5e086b4 · inbound

Analyzing and Mitigating Object Hallucination in Large Vision-Language Models cites this paper.

Analyzing and Mitigating Object Hallucination in Large Vision-Language Models Multi-Grained Vision Language Pre-Training: Aligning Texts with Visual Concepts

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-17T22:46:52.826787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T22:46:52.791128Z digest=sha256:40a296a97060f716c2ea93214dc264a05c72d2df2ed5c7a01c3141c12f34914f

Observation e43e2174-5234-4470-85c8-3e36a08f3c34 · inbound

The ART of Composition: Attention-Regularized Training for Compositional Visual Grounding cites this paper.

The ART of Composition: Attention-Regularized Training for Compositional Visual Grounding Multi-Grained Vision Language Pre-Training: Aligning Texts with Visual Concepts

Reference 59

Resolution
verified exact
arxiv_id, observed 2026-05-23T07:25:28.426307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-23T07:24:01.527093Z digest=sha256:9e0f568a9e5137b456cc5581f8a2e821ea49fc32781763be17c34fc762e5f708

Observation 647af6fe-30a2-4aef-9a4e-f3e83754c175 · inbound

Efficient Vision Language Model Fine-tuning for Text-based Person Anomaly Search cites this paper.

Efficient Vision Language Model Fine-tuning for Text-based Person Anomaly Search Multi-Grained Vision Language Pre-Training: Aligning Texts with Visual Concepts

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-09T05:30:25.558782Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T05:30:25.558782Z digest=sha256:586c9e8054ff9afd8e2eb7cb11b3b97d6659fd7ba86ad6032571395d7596c17a

Observation 80a0492a-2a43-446d-8752-4f308ffd5727 · inbound

Visual Agentic AI for Spatial Reasoning with a Dynamic API cites this paper.

Visual Agentic AI for Spatial Reasoning with a Dynamic API Multi-Grained Vision Language Pre-Training: Aligning Texts with Visual Concepts

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-08T14:23:01.386624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:23:01.386624Z digest=sha256:dfc8195c919ab80513cc378b3a35be0127d62476917ac7aeee88da0d7cd9eef7

Observation 46b16274-c468-4b0a-8b45-98c9f7579ca6 · inbound

Stitch-a-Demo: Video Demonstrations from Multistep Descriptions cites this paper.

Stitch-a-Demo: Video Demonstrations from Multistep Descriptions Multi-Grained Vision Language Pre-Training: Aligning Texts with Visual Concepts

Reference 74

Resolution
verified exact
arxiv_id, observed 2026-05-23T00:32:18.234866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-23T00:30:55.729900Z digest=sha256:8933a736d77a0e346f7e2b1acb351aecaa0951f100b0f3db9df712424005ca25

Observation ed6d5e29-6cf9-4f0b-bf71-73cdb6b0e7d3 · inbound

A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets cites this paper.

A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets Multi-Grained Vision Language Pre-Training: Aligning Texts with Visual Concepts

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T19:45:53.000363Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:45:53.000363Z digest=sha256:d08c9baf72acd944b2677072ad61b76f76baef34200aacc9ea1282600302f379

Observation e4f707dc-4835-4fa4-8c53-b371b38b35ad · inbound

Filter-And-Refine: A MLLM Based Cascade System for Industrial-Scale Video Content Moderation cites this paper.

Filter-And-Refine: A MLLM Based Cascade System for Industrial-Scale Video Content Moderation Multi-Grained Vision Language Pre-Training: Aligning Texts with Visual Concepts

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-06T14:58:44.002049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:58:44.002049Z digest=sha256:51f1ac2273fec2ab0cafc6b56356e59864185c04eb099c0bf6cb4cfeaf4a0d69

Observation ee10d92c-9187-42b2-82da-60aefc2d655d · inbound

WikiCLIP: An Efficient Contrastive Baseline for Open-domain Visual Entity Recognition cites this paper.

WikiCLIP: An Efficient Contrastive Baseline for Open-domain Visual Entity Recognition Multi-Grained Vision Language Pre-Training: Aligning Texts with Visual Concepts

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-15T13:15:50.590773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T13:11:54.384284Z digest=sha256:77f51c8ccec0577356e5eebbb51ceebf5fd8b271a8d3f5bafceef6fd7278c616

Observation cdab544e-40fb-40e6-8171-67f5202d82b3 · inbound

WikiCLIP: An Efficient Contrastive Baseline for Open-domain Visual Entity Recognition cites this paper.

WikiCLIP: An Efficient Contrastive Baseline for Open-domain Visual Entity Recognition Multi-Grained Vision Language Pre-Training: Aligning Texts with Visual Concepts

Reference 46

Resolution
unresolved
no resolver link, observed 2026-07-14T23:55:24.006436Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T23:55:24.006436Z digest=sha256:05b5642642e1a772fc2e6287ea2a4150ff6bc6f160af81181ba1fba0440a28ab

Observation f51790e1-a4d6-4a20-b16e-02358bcf9aae · inbound

InstAP: Instance-Aware Vision-Language Pre-Train for Spatial-Temporal Understanding cites this paper.

InstAP: Instance-Aware Vision-Language Pre-Train for Spatial-Temporal Understanding Multi-Grained Vision Language Pre-Training: Aligning Texts with Visual Concepts

Reference 67

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:30:58.271354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T18:06:33.139310Z digest=sha256:3e75af411a46ba17b42251d3f295992814d3426f77b999cd4fc4aefdf0e929e4

Observation af4bc09a-e5a3-42a5-ab58-80596b635f3c · inbound

Road Maps as Free Geometric Priors: Weather-Invariant Drone Geo-Localization with GeoFuse cites this paper.

Road Maps as Free Geometric Priors: Weather-Invariant Drone Geo-Localization with GeoFuse Multi-Grained Vision Language Pre-Training: Aligning Texts with Visual Concepts

Reference 49

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T20:55:04.310749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T20:49:31.140523Z digest=sha256:531d5d5597c629536829abf9c05854c5a461c27e1f5bf6141d95fc24037aaccb

Observation 9dbe92e3-e26e-4794-9155-2b41f3d913b6 · inbound

T-CLIP: Enabling Thermal Perception for Contrastive Language-Image Pretraining cites this paper.

T-CLIP: Enabling Thermal Perception for Contrastive Language-Image Pretraining Multi-Grained Vision Language Pre-Training: Aligning Texts with Visual Concepts

Reference 163

Resolution
metadata mismatch
arxiv_id, observed 2026-06-28T19:42:36.271287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-28T18:53:10.978631Z digest=sha256:fa07b39a8d299979d07eb5a8e12fa6e3de67392939c189d6d44f2f69fb80d057

Observation 87743c44-bb4a-476b-9e8f-a315a4d5ebd4 · inbound

Qwen-RobotWorld Technical Report: Unifying Embodied World Modeling through Language-Conditioned Video Generation cites this paper.

Qwen-RobotWorld Technical Report: Unifying Embodied World Modeling through Language-Conditioned Video Generation Multi-Grained Vision Language Pre-Training: Aligning Texts with Visual Concepts

Reference 153

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T17:18:43.861357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-27T04:19:26.332718Z digest=sha256:db9cd8227c82f66753cfe75dc0dd4cf3c665752a53e48321558340734ed8d1d6

Observation b6002f41-ad6d-4206-a94e-1e42ca5c00c2 · inbound

Qwen-Audio-VAE Technical Report cites this paper.

Qwen-Audio-VAE Technical Report Multi-Grained Vision Language Pre-Training: Aligning Texts with Visual Concepts

Reference 168

Resolution
unresolved
no resolver link, observed 2026-07-14T03:31:19.309532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T03:31:19.309532Z digest=sha256:1a3022eb18be01663aad81498e3ba05164a3128bad757763b01332f5e0043f85

Observation 7eb9a947-ea5e-4497-a68c-abe79cbc8ad0 · inbound

LAVIFT: Latent-Action-Guided Vision Fine-Tuning for Surgical Interaction Recognition cites this paper.

LAVIFT: Latent-Action-Guided Vision Fine-Tuning for Surgical Interaction Recognition Multi-Grained Vision Language Pre-Training: Aligning Texts with Visual Concepts

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-01T11:31:04.257952Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:31:04.257952Z digest=sha256:06e778df59aea19eb7be4da01cc897a658500e1bdad61bbcbe6cc503922f1a03

Observation 37206346-f298-4f26-8e95-aa3023989e37 · inbound

DICA: Dual-Indicator Guided Contrastive Alignment in Multimodal Large Language Models cites this paper.

DICA: Dual-Indicator Guided Contrastive Alignment in Multimodal Large Language Models Multi-Grained Vision Language Pre-Training: Aligning Texts with Visual Concepts

Reference 62

Resolution
unresolved
no resolver link, observed 2026-07-31T23:32:09.979514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T23:32:09.979514Z digest=sha256:e71dc9b39be6a43dd2255ce6dfbf59b1668b81c71dc07a0fbe7392ee6faa7c8c