Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 17 inbound Pith citation observations for arXiv:2111.08276.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-09T05:30:25.558782Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-03T17:18:43.859907Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation 81367072-436f-4468-a3ee-66063c46e2dc · inbound
ViperGPT: Visual Inference via Python Execution for Reasoning Multi-Grained Vision Language Pre-Training: Aligning Texts with Visual Concepts
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 1f50a754-badd-479f-98f7-5b14e5e086b4 · inbound
Analyzing and Mitigating Object Hallucination in Large Vision-Language Models Multi-Grained Vision Language Pre-Training: Aligning Texts with Visual Concepts
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation e43e2174-5234-4470-85c8-3e36a08f3c34 · inbound
The ART of Composition: Attention-Regularized Training for Compositional Visual Grounding Multi-Grained Vision Language Pre-Training: Aligning Texts with Visual Concepts
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 647af6fe-30a2-4aef-9a4e-f3e83754c175 · inbound
Efficient Vision Language Model Fine-tuning for Text-based Person Anomaly Search Multi-Grained Vision Language Pre-Training: Aligning Texts with Visual Concepts
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 80a0492a-2a43-446d-8752-4f308ffd5727 · inbound
Visual Agentic AI for Spatial Reasoning with a Dynamic API Multi-Grained Vision Language Pre-Training: Aligning Texts with Visual Concepts
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 46b16274-c468-4b0a-8b45-98c9f7579ca6 · inbound
Stitch-a-Demo: Video Demonstrations from Multistep Descriptions Multi-Grained Vision Language Pre-Training: Aligning Texts with Visual Concepts
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation ed6d5e29-6cf9-4f0b-bf71-73cdb6b0e7d3 · inbound
A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets Multi-Grained Vision Language Pre-Training: Aligning Texts with Visual Concepts
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e4f707dc-4835-4fa4-8c53-b371b38b35ad · inbound
Filter-And-Refine: A MLLM Based Cascade System for Industrial-Scale Video Content Moderation Multi-Grained Vision Language Pre-Training: Aligning Texts with Visual Concepts
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ee10d92c-9187-42b2-82da-60aefc2d655d · inbound
WikiCLIP: An Efficient Contrastive Baseline for Open-domain Visual Entity Recognition Multi-Grained Vision Language Pre-Training: Aligning Texts with Visual Concepts
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation cdab544e-40fb-40e6-8171-67f5202d82b3 · inbound
WikiCLIP: An Efficient Contrastive Baseline for Open-domain Visual Entity Recognition Multi-Grained Vision Language Pre-Training: Aligning Texts with Visual Concepts
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f51790e1-a4d6-4a20-b16e-02358bcf9aae · inbound
InstAP: Instance-Aware Vision-Language Pre-Train for Spatial-Temporal Understanding Multi-Grained Vision Language Pre-Training: Aligning Texts with Visual Concepts
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation af4bc09a-e5a3-42a5-ab58-80596b635f3c · inbound
Road Maps as Free Geometric Priors: Weather-Invariant Drone Geo-Localization with GeoFuse Multi-Grained Vision Language Pre-Training: Aligning Texts with Visual Concepts
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 9dbe92e3-e26e-4794-9155-2b41f3d913b6 · inbound
T-CLIP: Enabling Thermal Perception for Contrastive Language-Image Pretraining Multi-Grained Vision Language Pre-Training: Aligning Texts with Visual Concepts
Reference 163
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 87743c44-bb4a-476b-9e8f-a315a4d5ebd4 · inbound
Qwen-RobotWorld Technical Report: Unifying Embodied World Modeling through Language-Conditioned Video Generation Multi-Grained Vision Language Pre-Training: Aligning Texts with Visual Concepts
Reference 153
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation b6002f41-ad6d-4206-a94e-1e42ca5c00c2 · inbound
Qwen-Audio-VAE Technical Report Multi-Grained Vision Language Pre-Training: Aligning Texts with Visual Concepts
Reference 168
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7eb9a947-ea5e-4497-a68c-abe79cbc8ad0 · inbound
LAVIFT: Latent-Action-Guided Vision Fine-Tuning for Surgical Interaction Recognition Multi-Grained Vision Language Pre-Training: Aligning Texts with Visual Concepts
Reference 83
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 37206346-f298-4f26-8e95-aa3023989e37 · inbound
DICA: Dual-Indicator Guided Contrastive Alignment in Multimodal Large Language Models Multi-Grained Vision Language Pre-Training: Aligning Texts with Visual Concepts
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.