Pith. sign in

Paper Citation Record · LEDGER

Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 33 inbound Pith citation observations for arXiv:2102.05918.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2102.05918 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 33 of 33 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 33 of 33 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T22:41:30.991417Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

1196
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation eb4228ba-fe36-47c9-87b6-027482e7a144 · inbound

LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs cites this paper.

LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:21:01.079608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-12T10:21:01.062199Z digest=sha256:1d9259d3568bd5a9246c16fe28c5a2d10395fab6a1ebe1ab193cef1a6ee67ca3

Observation 6085eebd-ed68-41ce-b8dc-6c865a7e665b · inbound

Florence: A New Foundation Model for Computer Vision cites this paper.

Florence: A New Foundation Model for Computer Vision Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-16T09:38:09.518303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-16T09:38:09.427509Z digest=sha256:4d2b21cc099786d9089ba617ee423b254b92a9ac98cd3678a71c0331c18e4d74

Observation 7fd6cab6-82f1-4d96-a6d7-6e0c3bc84eb8 · inbound

Flamingo: a Visual Language Model for Few-Shot Learning cites this paper.

Flamingo: a Visual Language Model for Few-Shot Learning Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-12T04:22:30.583825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-12T04:22:30.008355Z digest=sha256:07c84ba8aece521e47b7d38e83531221500f97c21f9bbe23610d95b46cdaf62d

Observation 2d8c5e76-6c40-44b1-b0b6-48f810cc2611 · inbound

DetailCLIP: Injecting Image Details into CLIP's Feature Space cites this paper.

DetailCLIP: Injecting Image Details into CLIP's Feature Space Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-24T11:09:22.398768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-24T11:08:20.298043Z digest=sha256:75159432af20177a37f6857f4ce59f5d7eadd913bb178173d549d9d436cba4cd

Observation f0e83b57-fdfe-46bf-9ec6-98f03ef143af · inbound

LAION-5B: An open large-scale dataset for training next generation image-text models cites this paper.

LAION-5B: An open large-scale dataset for training next generation image-text models Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-13T14:22:17.247569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-13T14:22:16.968028Z digest=sha256:acd3d22ba54636f1a4935d63488fc17b64f1bd7c2bf97bdad22fd53224d85b62

Observation 395f5349-5278-4d09-b17d-16c4222c7f42 · inbound

Editing Models with Task Arithmetic cites this paper.

Editing Models with Task Arithmetic Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision

Reference 42

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T08:09:13.018298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-13T08:09:12.716163Z digest=sha256:3bf96384bb4cde1d020bffdc14b4a73cf9628120ceb9670175515543696e4ce2

Observation 7e9cc680-71f9-4ce0-9e53-16054243c38d · inbound

BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models cites this paper.

BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-12T00:10:49.486795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-12T00:10:48.610351Z digest=sha256:d35fe7846f94d5f1e6fbb809f85323e8a0a3e62367d1f765dcdfda887f3f05cd

Observation f433eb90-0de5-46e5-b917-1700b2cb27ab · inbound

OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models cites this paper.

OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T09:55:35.565340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-17T09:55:35.452649Z digest=sha256:8a44f22b1544ec143695822ec466adb6805c4d1a4d0609d1d0cd5af8d8a9b4e9

Observation c3fd3fae-5c19-434f-930b-fbb8d099ae27 · inbound

Color in Visual-Language Models: CLIP deficiencies cites this paper.

Color in Visual-Language Models: CLIP deficiencies Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-08T22:41:30.991417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T22:41:30.991417Z digest=sha256:166d293826b00a94f6a890447f9f53e075ca6359d673592b90f213bf09112ab0

Observation a180953a-5e8a-4335-8ac2-2759aeb5e62c · inbound

A Real-to-Sim-to-Real Approach to Robotic Manipulation with VLM-Generated Iterative Keypoint Rewards cites this paper.

A Real-to-Sim-to-Real Approach to Robotic Manipulation with VLM-Generated Iterative Keypoint Rewards Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-08T00:03:16.291602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T00:03:16.291602Z digest=sha256:8d43876fb34a6cf914c8aec11d756a0abee03cea78d197717f277547fd8d8a45

Observation 1ffc0213-69d9-479c-9df5-8ce397816749 · inbound

A Survey on Training-free Open-Vocabulary Semantic Segmentation cites this paper.

A Survey on Training-free Open-Vocabulary Semantic Segmentation Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T13:15:35.615457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:15:35.615457Z digest=sha256:b88817d443ff4ad026442b4670e34e00d85f31719fd12480b907ed748a8b9018

Observation f51deee5-ada3-4f9b-867c-8f0d56956134 · inbound

WisWheat: A Three-Tiered Vision-Language Dataset for Wheat Management cites this paper.

WisWheat: A Three-Tiered Vision-Language Dataset for Wheat Management Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:10.304185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:02:10.304185Z digest=sha256:f7f1932c6871c514d92355a3352bb7e0dfbbfbe93bba8d9c3d0f96c29c642a9c

Observation c4c8289e-bd1a-4e81-8b09-344fef440dec · inbound

Visual Pre-Training on Unlabeled Images using Reinforcement Learning cites this paper.

Visual Pre-Training on Unlabeled Images using Reinforcement Learning Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T01:06:19.262041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T01:06:19.262041Z digest=sha256:8797f51452c399bbeb9bfa35d0c76067694fa19d0086980fe63cd44685a79b84

Observation 02c019f0-7f4e-4b7e-b6b5-d13924477716 · inbound

CF-VLM:CounterFactual Vision-Language Fine-tuning cites this paper.

CF-VLM:CounterFactual Vision-Language Fine-tuning Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:07.069602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:07.069602Z digest=sha256:27bb74ca32bf720fe1daa5e0c02073bc1a212b3625100e99351f8a28fe3d9428

Observation 69f0053e-9912-4700-8fda-4e1262126377 · inbound

PaCo-FR: Patch-Pixel Aligned End-to-End Codebook Learning for Facial Representation Pre-training cites this paper.

PaCo-FR: Patch-Pixel Aligned End-to-End Codebook Learning for Facial Representation Pre-training Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision

Reference 29

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T22:56:53.171154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-18T22:53:23.271423Z digest=sha256:03564cbcc79282d3571a900ae237ac101b8b41b8d0621d6f7ca61add5b9a11cb

Observation b07fa155-e0df-4e0b-9d51-b6569f4a917e · inbound

Robust and Label-Efficient Deep Waste Detection cites this paper.

Robust and Label-Efficient Deep Waste Detection Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T16:15:06.970353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:15:06.970353Z digest=sha256:97c9a33dbc1cb9151b000eccb2e66992e17a9ff932eac749f1fe96caff9303d4

Observation 2504c725-81df-4bb4-873e-cf83c4670f93 · inbound

VLMs-in-the-Wild: Bridging the Gap Between Academic Benchmarks and Enterprise Reality cites this paper.

VLMs-in-the-Wild: Bridging the Gap Between Academic Benchmarks and Enterprise Reality Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-05T11:11:20.762326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:11:20.762326Z digest=sha256:cdd383388178ff877ea6e20d5ecd465de99e453d8bea6f87ad58b9de5ba35b84

Observation 0f60d7bd-3257-4cc5-bc51-5fb93d32a14c · inbound

Rate-Distortion Limits for Multimodal Retrieval: Theory, Optimal Codes, and Finite-Sample Guarantees cites this paper.

Rate-Distortion Limits for Multimodal Retrieval: Theory, Optimal Codes, and Finite-Sample Guarantees Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-04T17:18:53.610342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T17:18:53.610342Z digest=sha256:b4125d7b31dc594a6ee7ced7c59815b2ad1e4deedcecc2d0581f3a6eb8ba772b

Observation 8fa62d87-4c65-4139-9306-84d709e30d4f · inbound

The Vision Wormhole: Latent-Space Communication in Heterogeneous Multi-Agent Systems cites this paper.

The Vision Wormhole: Latent-Space Communication in Heterogeneous Multi-Agent Systems Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-02T22:58:07.375228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T22:58:07.375228Z digest=sha256:93764632bfd11ffd972120a17e11c6222cce780f14d996f78e4feb3ee4a130ba

Observation 79e88e60-3b7d-44f2-a561-87a2df459541 · inbound

WikiCLIP: An Efficient Contrastive Baseline for Open-domain Visual Entity Recognition cites this paper.

WikiCLIP: An Efficient Contrastive Baseline for Open-domain Visual Entity Recognition Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-15T13:15:50.553661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T13:11:54.384284Z digest=sha256:2a6c00fc271aa43a1d62e1d194092f08960e7881433f82dd8380cba1d2c81f5e

Observation 456e6001-3639-4c05-8ea9-b06b04fc89fe · inbound

WikiCLIP: An Efficient Contrastive Baseline for Open-domain Visual Entity Recognition cites this paper.

WikiCLIP: An Efficient Contrastive Baseline for Open-domain Visual Entity Recognition Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-14T23:55:24.006436Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T23:55:24.006436Z digest=sha256:112395e4841b8ad452508de2b17e8835564c9e98ff546805defb08fc4559fc32

Observation be20fb00-f225-4467-92b9-931a212c9bb9 · inbound

Explicit Logic Channel for Validation and Enhancement of MLLMs on Zero-Shot Tasks cites this paper.

Explicit Logic Channel for Validation and Enhancement of MLLMs on Zero-Shot Tasks Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-21T11:10:02.024909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-21T11:09:51.816554Z digest=sha256:80e26ade81a4d648c443ae85e4610bd843f2fab145422ff1f5064621c3c3a995

Observation 3d153cec-cc76-4407-83ec-879e1681f651 · inbound

DeCo-DETR: Decoupled Cognition DETR for efficient Open-Vocabulary Object Detection cites this paper.

DeCo-DETR: Decoupled Cognition DETR for efficient Open-Vocabulary Object Detection Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T20:33:16.707134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-13T20:30:16.163179Z digest=sha256:765eafafa4c1e7a0ebf5136e98c51d2ac5466cdb0e27bfdbbf0d5bd05641b785

Observation a5243118-073a-4d3d-9ff6-355111f88bfb · inbound

DeCo-DETR: Decoupled Cognition DETR for efficient Open-Vocabulary Object Detection cites this paper.

DeCo-DETR: Decoupled Cognition DETR for efficient Open-Vocabulary Object Detection Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T21:49:29.323542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-14T21:48:49.892964Z digest=sha256:f52f1c48d61924df957eebd547bae9327c7b72ed97464f4af5fa9d25ba011751

Observation 2eac7ce6-23f4-4ea8-ba63-a6293d0feb53 · inbound

Latent Anomaly Knowledge Excavation: Unveiling Sparse Sensitive Neurons in Vision-Language Models cites this paper.

Latent Anomaly Knowledge Excavation: Unveiling Sparse Sensitive Neurons in Vision-Language Models Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T06:51:16.216373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T17:25:34.942028Z digest=sha256:4d069c8f104ddbc035ed6b65c767fb6b99b22bde458e2839ae91ed0c35f7d679

Observation 4e82d205-e315-41a4-97c8-1d0208ad0f87 · inbound

Compared to What? Baselines and Metrics for Counterfactual Prompting cites this paper.

Compared to What? Baselines and Metrics for Counterfactual Prompting Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision

Reference 93

Resolution
verified exact
arxiv_id, observed 2026-05-09T19:05:10.548305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-09T19:02:46.991897Z digest=sha256:4fce3be5ee2800abc77e8b768696c4e3fe77790669a29aa81f6f112ca5311620

Observation ef87c4ca-e90d-4d07-a1d1-47f52356d4dc · inbound

Vision Harnessing Agent for Open Ad-hoc Segmentation cites this paper.

Vision Harnessing Agent for Open Ad-hoc Segmentation Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-20T05:53:04.439800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-20T05:52:40.429412Z digest=sha256:de5be45a4215c4b71b48eeb8a4beee09b35d4d4de99402a7322a9182bb32acb1

Observation 44908e21-1cef-439e-b18d-57458df842c3 · inbound

Toward Calibrated, Fair, and accurate Deepfake Detection cites this paper.

Toward Calibrated, Fair, and accurate Deepfake Detection Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision

Reference 266

Resolution
verified exact
arxiv_id, observed 2026-06-28T07:11:45.290744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-28T07:05:18.026601Z digest=sha256:e41710e18c463cd4aad7609bfdf5a84121a29c0a6c2ac88b169887a72fcd5e80

Observation 11c0e04f-f341-4fa2-9ced-cdeb5eab031b · inbound

Modeling Complex Behaviors: Multi-Personality Composition and Dynamic Switching in Vision-Language Models cites this paper.

Modeling Complex Behaviors: Multi-Personality Composition and Dynamic Switching in Vision-Language Models Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision

Reference 74

Resolution
verified exact
arxiv_id, observed 2026-07-03T05:47:41.829444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-27T13:04:14.886733Z digest=sha256:bda5df12df91a89cf488d266e99a5ea396366a7597ee8b236205b4a6c0fed2b9

Observation 887c66e1-1378-4117-a8ea-35b4aad639d3 · inbound

Qwen-RobotWorld Technical Report: Unifying Embodied World Modeling through Language-Conditioned Video Generation cites this paper.

Qwen-RobotWorld Technical Report: Unifying Embodied World Modeling through Language-Conditioned Video Generation Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision

Reference 140

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T17:18:43.932557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-27T04:19:26.332718Z digest=sha256:fd4afb7e2a275b066f7394b3dfa50251686def7e095b1bfe4edfeb3e1b429826

Observation daa1c166-07ec-40c9-b3f6-4c193ffb0d88 · inbound

Combating Textual Noise and Redundancy: Entropy-Aware Dense Visual Token Pruning cites this paper.

Combating Textual Noise and Redundancy: Entropy-Aware Dense Visual Token Pruning Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision

Reference 22

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T14:48:32.368383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-03T14:47:35.377391Z digest=sha256:65bcb8ce1806e4da061521716f2a0af06e2f5dfd1f5fee492a999f073efe95a9

Observation 2a4df5d4-58ff-4181-a3bd-baefb693cf78 · inbound

Qwen-Audio-VAE Technical Report cites this paper.

Qwen-Audio-VAE Technical Report Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision

Reference 155

Resolution
unresolved
no resolver link, observed 2026-07-14T03:31:19.309532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T03:31:19.309532Z digest=sha256:22fe15e8c592865d88de315e861e1fb270db14bb354b96692a8cc385b65d34a7

Observation d4b3338b-8b46-4a9c-a9ae-6a354c1cdbcc · inbound

Theia: Large-Scale Multimodal Captioning and Automated Validation of the Incidents1M Dataset for Data-Free Distillation cites this paper.

Theia: Large-Scale Multimodal Captioning and Automated Validation of the Incidents1M Dataset for Data-Free Distillation Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-31T12:44:13.340533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T12:44:13.340533Z digest=sha256:9051c162b38ee6202794e1fd76506af8bae7c67fffe1ed96b1ab63a6f2613942