Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 33 inbound Pith citation observations for arXiv:2102.05918.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-08T22:41:30.991417Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z
0 of 0 outbound references displayed
External citation measurements
1196
arxiv_reference, observed 2026-08-05T02:28:24.338817Z
No outbound reference observations are available for this paper version.
Observation eb4228ba-fe36-47c9-87b6-027482e7a144 · inbound
LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 6085eebd-ed68-41ce-b8dc-6c865a7e665b · inbound
Florence: A New Foundation Model for Computer Vision Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 7fd6cab6-82f1-4d96-a6d7-6e0c3bc84eb8 · inbound
Flamingo: a Visual Language Model for Few-Shot Learning Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 2d8c5e76-6c40-44b1-b0b6-48f810cc2611 · inbound
DetailCLIP: Injecting Image Details into CLIP's Feature Space Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f0e83b57-fdfe-46bf-9ec6-98f03ef143af · inbound
LAION-5B: An open large-scale dataset for training next generation image-text models Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 395f5349-5278-4d09-b17d-16c4222c7f42 · inbound
Editing Models with Task Arithmetic Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 7e9cc680-71f9-4ce0-9e53-16054243c38d · inbound
BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f433eb90-0de5-46e5-b917-1700b2cb27ab · inbound
OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation c3fd3fae-5c19-434f-930b-fbb8d099ae27 · inbound
Color in Visual-Language Models: CLIP deficiencies Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a180953a-5e8a-4335-8ac2-2759aeb5e62c · inbound
A Real-to-Sim-to-Real Approach to Robotic Manipulation with VLM-Generated Iterative Keypoint Rewards Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1ffc0213-69d9-479c-9df5-8ce397816749 · inbound
A Survey on Training-free Open-Vocabulary Semantic Segmentation Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f51deee5-ada3-4f9b-867c-8f0d56956134 · inbound
WisWheat: A Three-Tiered Vision-Language Dataset for Wheat Management Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c4c8289e-bd1a-4e81-8b09-344fef440dec · inbound
Visual Pre-Training on Unlabeled Images using Reinforcement Learning Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 02c019f0-7f4e-4b7e-b6b5-d13924477716 · inbound
CF-VLM:CounterFactual Vision-Language Fine-tuning Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 69f0053e-9912-4700-8fda-4e1262126377 · inbound
PaCo-FR: Patch-Pixel Aligned End-to-End Codebook Learning for Facial Representation Pre-training Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation b07fa155-e0df-4e0b-9d51-b6569f4a917e · inbound
Robust and Label-Efficient Deep Waste Detection Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2504c725-81df-4bb4-873e-cf83c4670f93 · inbound
VLMs-in-the-Wild: Bridging the Gap Between Academic Benchmarks and Enterprise Reality Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0f60d7bd-3257-4cc5-bc51-5fb93d32a14c · inbound
Rate-Distortion Limits for Multimodal Retrieval: Theory, Optimal Codes, and Finite-Sample Guarantees Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8fa62d87-4c65-4139-9306-84d709e30d4f · inbound
The Vision Wormhole: Latent-Space Communication in Heterogeneous Multi-Agent Systems Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 79e88e60-3b7d-44f2-a561-87a2df459541 · inbound
WikiCLIP: An Efficient Contrastive Baseline for Open-domain Visual Entity Recognition Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 456e6001-3639-4c05-8ea9-b06b04fc89fe · inbound
WikiCLIP: An Efficient Contrastive Baseline for Open-domain Visual Entity Recognition Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation be20fb00-f225-4467-92b9-931a212c9bb9 · inbound
Explicit Logic Channel for Validation and Enhancement of MLLMs on Zero-Shot Tasks Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 3d153cec-cc76-4407-83ec-879e1681f651 · inbound
DeCo-DETR: Decoupled Cognition DETR for efficient Open-Vocabulary Object Detection Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a5243118-073a-4d3d-9ff6-355111f88bfb · inbound
DeCo-DETR: Decoupled Cognition DETR for efficient Open-Vocabulary Object Detection Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 2eac7ce6-23f4-4ea8-ba63-a6293d0feb53 · inbound
Latent Anomaly Knowledge Excavation: Unveiling Sparse Sensitive Neurons in Vision-Language Models Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 4e82d205-e315-41a4-97c8-1d0208ad0f87 · inbound
Compared to What? Baselines and Metrics for Counterfactual Prompting Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision
Reference 93
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation ef87c4ca-e90d-4d07-a1d1-47f52356d4dc · inbound
Vision Harnessing Agent for Open Ad-hoc Segmentation Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 44908e21-1cef-439e-b18d-57458df842c3 · inbound
Toward Calibrated, Fair, and accurate Deepfake Detection Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision
Reference 266
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 11c0e04f-f341-4fa2-9ced-cdeb5eab031b · inbound
Modeling Complex Behaviors: Multi-Personality Composition and Dynamic Switching in Vision-Language Models Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 887c66e1-1378-4117-a8ea-35b4aad639d3 · inbound
Qwen-RobotWorld Technical Report: Unifying Embodied World Modeling through Language-Conditioned Video Generation Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision
Reference 140
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation daa1c166-07ec-40c9-b3f6-4c193ffb0d88 · inbound
Combating Textual Noise and Redundancy: Entropy-Aware Dense Visual Token Pruning Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 2a4df5d4-58ff-4181-a3bd-baefb693cf78 · inbound
Qwen-Audio-VAE Technical Report Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision
Reference 155
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d4b3338b-8b46-4a9c-a9ae-6a354c1cdbcc · inbound
Theia: Large-Scale Multimodal Captioning and Automated Validation of the Incidents1M Dataset for Data-Free Distillation Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.