Pith. sign in

Paper Citation Record · LEDGER

Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision

As of 22 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 50 inbound Pith citation observations for arXiv:2102.05918.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2102.05918 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 50 of 50 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 50 of 50 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T00:12:08.581177Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

1196
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation eb4228ba-fe36-47c9-87b6-027482e7a144 · inbound

LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs cites this paper.

LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:21:01.079608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-12T10:21:01.062199Z digest=sha256:890df59da4942ac95cc812f0007eb2f8bc85d0208a3693b5ceeb8a5d9a4bc2f4

Observation 6085eebd-ed68-41ce-b8dc-6c865a7e665b · inbound

Florence: A New Foundation Model for Computer Vision cites this paper.

Florence: A New Foundation Model for Computer Vision Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-16T09:38:09.518303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-16T09:38:09.427509Z digest=sha256:26e51a52b0da4372aad450e060e0f926e213e66090951ec6df8dcb6756e14909

Observation 7fd6cab6-82f1-4d96-a6d7-6e0c3bc84eb8 · inbound

Flamingo: a Visual Language Model for Few-Shot Learning cites this paper.

Flamingo: a Visual Language Model for Few-Shot Learning Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-12T04:22:30.583825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-12T04:22:30.008355Z digest=sha256:e2b7d28b7acee5a415be182f7de7c07d680cdee37b6238c3eb5a5b4f05cfdb29

Observation 2d8c5e76-6c40-44b1-b0b6-48f810cc2611 · inbound

DetailCLIP: Injecting Image Details into CLIP's Feature Space cites this paper.

DetailCLIP: Injecting Image Details into CLIP's Feature Space Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-24T11:09:22.398768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-24T11:08:20.298043Z digest=sha256:48813b2d5645fb80e963f6023e38721fac2f1719ae233c07c443009b4701d77f

Observation f0e83b57-fdfe-46bf-9ec6-98f03ef143af · inbound

LAION-5B: An open large-scale dataset for training next generation image-text models cites this paper.

LAION-5B: An open large-scale dataset for training next generation image-text models Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-13T14:22:17.247569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-13T14:22:16.968028Z digest=sha256:00489dbf8913471cc4cb394bf373e6d356c18355c6fb45bd018c7984b5740800

Observation 395f5349-5278-4d09-b17d-16c4222c7f42 · inbound

Editing Models with Task Arithmetic cites this paper.

Editing Models with Task Arithmetic Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision

Reference 42

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T08:09:13.018298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-13T08:09:12.716163Z digest=sha256:42625557e0b8d23309830e5e06b20a4a5a4f0ba6bba62fa94693b70ef41e39d8

Observation 7e9cc680-71f9-4ce0-9e53-16054243c38d · inbound

BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models cites this paper.

BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-12T00:10:49.486795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-12T00:10:48.610351Z digest=sha256:b5baac0e387a62c45eb733b9ddaf8b6f79fdb1a752323f9661565e30074f19b4

Observation f433eb90-0de5-46e5-b917-1700b2cb27ab · inbound

OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models cites this paper.

OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T09:55:35.565340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-17T09:55:35.452649Z digest=sha256:a0129449d4adc1abf9324df04df0bf58df4d22228411646192eee9b6c39b8d14

Observation f58c1641-f862-48e3-b6c6-3f1f598e086f · inbound

MuLan: Adapting Multilingual Diffusion Models for Hundreds of Languages with Negligible Cost cites this paper.

MuLan: Adapting Multilingual Diffusion Models for Hundreds of Languages with Negligible Cost Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T04:34:52.512550Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:34:52.512550Z digest=sha256:5f0e177bed8f637ecc93cc0470238416a80a167046a4db7ab09e014f52515656

Observation 6e4539fe-dfb6-4e89-bea0-28536504ee53 · inbound

Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension cites this paper.

Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T22:18:01.462379Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:18:01.462379Z digest=sha256:ecdd9351a076a0f83f7c122db3e28cdf000cb829dd5544869ee0b0999d508fd1

Observation 9a703ba8-af43-4ffd-80e8-072317af3141 · inbound

RefSAM3D: Adapting SAM with Cross-modal Reference for 3D Medical Image Segmentation cites this paper.

RefSAM3D: Adapting SAM with Cross-modal Reference for 3D Medical Image Segmentation Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T20:39:32.454362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:39:32.454362Z digest=sha256:162d6ef6407f081a1e999132e4532a247d33db25a1234dbd1f513b4bb6690c82

Observation 1d9c4445-c613-430a-8dc7-8b064178e0b1 · inbound

Foundation Models and Adaptive Feature Selection: A Synergistic Approach to Video Question Answering cites this paper.

Foundation Models and Adaptive Feature Selection: A Synergistic Approach to Video Question Answering Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T17:16:12.284203Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:16:12.284203Z digest=sha256:01f599bba94732989dc474e5b32ea634f5b8778c97ae7a9f6963b87e7c31ae52

Observation 421d0591-c9c0-4f80-bda5-516d35888bb2 · inbound

Does VLM Classification Benefit from LLM Description Semantics? cites this paper.

Does VLM Classification Benefit from LLM Description Semantics? Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T14:31:37.223563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:31:37.223563Z digest=sha256:84c22497d108e5d95b445f4bda2c0d6bb1ce0d19ffc4d83d1f0d062d8dbff109

Observation 51ccd206-bab0-4b0f-9913-00968d7d84b7 · inbound

Visualizing the Invisible: A Generative AR System for Intuitive Multi-Modal Sensor Data Presentation cites this paper.

Visualizing the Invisible: A Generative AR System for Intuitive Multi-Modal Sensor Data Presentation Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-11T13:08:48.345454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:08:48.345454Z digest=sha256:c19266d46de38a5c4560d988d9620e775c676e58cfabf0224ac944b24ef88050

Observation d4e31344-cc47-4f90-869d-a5e383d9a828 · inbound

A Decade of Deep Learning: A Survey on The Magnificent Seven cites this paper.

A Decade of Deep Learning: A Survey on The Magnificent Seven Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T16:10:24.421103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:10:24.421103Z digest=sha256:e590793cbe0a28304aea28b46fe7e5edac317af9937c0afd484f16a37723b053

Observation cdcd0b62-1ac9-42fd-beca-718e10dd8200 · inbound

ViPCap: Retrieval Text-Based Visual Prompts for Lightweight Image Captioning cites this paper.

ViPCap: Retrieval Text-Based Visual Prompts for Lightweight Image Captioning Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T00:47:52.398127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:47:52.398127Z digest=sha256:c5e375d2a2ab8d98a11aecceb787c9f1ecb0a53676e8830347cd8147040110f5

Observation 7869bfb3-e5ba-4c1e-9e0a-18871455f695 · inbound

Probing Visual Language Priors in VLMs cites this paper.

Probing Visual Language Priors in VLMs Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-10T22:55:53.964888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:55:53.964888Z digest=sha256:a4fdc60467ecd25609b70862b596e4c891dec7bed9dd6c31cab216608f522f1c

Observation f46d9fff-300a-4c40-8208-ad430b63b1f4 · inbound

Efficient Domain Adaptation of Multimodal Embeddings using Constrastive Learning cites this paper.

Efficient Domain Adaptation of Multimodal Embeddings using Constrastive Learning Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-09T13:37:11.808123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T13:37:11.808123Z digest=sha256:3600751e6238e7abdfded5570b81cf2b48badbba4ee885bac1bcbbe3e47b8ae8

Observation 2598b6c4-70af-436b-a084-1e18ef239996 · inbound

Mirai: A Wearable Proactive AI "Inner-Voice" for Contextual Nudging cites this paper.

Mirai: A Wearable Proactive AI "Inner-Voice" for Contextual Nudging Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-09T12:25:08.797215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T12:25:08.797215Z digest=sha256:8f1f34c40d06490767c45dbc4744b988212fb500fcaa205102d4e842fdf71349

Observation c3fd3fae-5c19-434f-930b-fbb8d099ae27 · inbound

Color in Visual-Language Models: CLIP deficiencies cites this paper.

Color in Visual-Language Models: CLIP deficiencies Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-08T22:41:30.991417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T22:41:30.991417Z digest=sha256:c105cb34f680c566be45f3caff795704f46096aab4c99fbf70d972234557306c

Observation a180953a-5e8a-4335-8ac2-2759aeb5e62c · inbound

A Real-to-Sim-to-Real Approach to Robotic Manipulation with VLM-Generated Iterative Keypoint Rewards cites this paper.

A Real-to-Sim-to-Real Approach to Robotic Manipulation with VLM-Generated Iterative Keypoint Rewards Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-08T00:03:16.291602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T00:03:16.291602Z digest=sha256:9c625dc4ba8160b67c4a8ddedf1ba52671c3fb688b6087c76709dcc683470046

Observation 5cfe1e3e-0b47-40c5-be01-dc4a7d0c8ca6 · inbound

Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models cites this paper.

Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision

Reference 189

Resolution
unresolved
no resolver link, observed 2026-08-15T23:21:12.793288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:21:12.793288Z digest=sha256:95774cd95f65d08b6f09f3a8273d7ece3f14a21dc7d1913ec43cde050df674f0

Observation 1e961784-7c43-4a20-a059-1805906fdc6f · inbound

Computer Vision Models Show Human-Like Sensitivity to Geometric and Topological Concepts cites this paper.

Computer Vision Models Show Human-Like Sensitivity to Geometric and Topological Concepts Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:48.869185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:21:48.869185Z digest=sha256:94440da8004107dd6dc505b8989e8f14f79c70f1544d9d3d367afd1ff4f041b4

Observation 1ffc0213-69d9-479c-9df5-8ce397816749 · inbound

A Survey on Training-free Open-Vocabulary Semantic Segmentation cites this paper.

A Survey on Training-free Open-Vocabulary Semantic Segmentation Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T13:15:35.615457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:15:35.615457Z digest=sha256:5da4f35ff4075254e2503ce36260196a17787bbf5299415f07076a00a6a676ba

Observation f51deee5-ada3-4f9b-867c-8f0d56956134 · inbound

WisWheat: A Three-Tiered Vision-Language Dataset for Wheat Management cites this paper.

WisWheat: A Three-Tiered Vision-Language Dataset for Wheat Management Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:10.304185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:02:10.304185Z digest=sha256:c1ee9aa6760dcb39f1a2272ba6b6c7356ffd8e8c1d22a6e9a187c1e7c74f1925

Observation c4c8289e-bd1a-4e81-8b09-344fef440dec · inbound

Visual Pre-Training on Unlabeled Images using Reinforcement Learning cites this paper.

Visual Pre-Training on Unlabeled Images using Reinforcement Learning Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T01:06:19.262041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T01:06:19.262041Z digest=sha256:1d038294ab11065e56ac57cbdf2d7b0e128f453aef4c75662975729596dac6cf

Observation 02c019f0-7f4e-4b7e-b6b5-d13924477716 · inbound

CF-VLM:CounterFactual Vision-Language Fine-tuning cites this paper.

CF-VLM:CounterFactual Vision-Language Fine-tuning Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:07.069602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:07.069602Z digest=sha256:4a262498f25f8bb9fc7329518cd0607094a62a28199e2571138b762d44cfdff0

Observation 6b9a24a1-bcfe-45ed-99ca-0c0b7e41e17d · inbound

AME: Aligned Manifold Entropy for Robust Vision-Language Distillation cites this paper.

AME: Aligned Manifold Entropy for Robust Vision-Language Distillation Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T17:42:20.457731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:42:20.457731Z digest=sha256:a921e0e0654561083142bf0c4e5f0f3a50de5fd0a218165a58150bba5e53034f

Observation 69f0053e-9912-4700-8fda-4e1262126377 · inbound

PaCo-FR: Patch-Pixel Aligned End-to-End Codebook Learning for Facial Representation Pre-training cites this paper.

PaCo-FR: Patch-Pixel Aligned End-to-End Codebook Learning for Facial Representation Pre-training Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision

Reference 29

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T22:56:53.171154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-18T22:53:23.271423Z digest=sha256:e364ce2498de58e25208bf1b68b0a7104ab850ff6b3fd12cd15b4bc6b79540e2

Observation 09f374f6-ef52-41bd-ba6e-19edde942969 · inbound

UniECS: Unified Multimodal E-Commerce Search Framework with Gated Cross-modal Fusion cites this paper.

UniECS: Unified Multimodal E-Commerce Search Framework with Gated Cross-modal Fusion Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T17:15:20.695184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:15:20.695184Z digest=sha256:c147316b07361cfe9a65915d8a00d81d751fad76b1fd32f9a154c3142d4b2a53

Observation b07fa155-e0df-4e0b-9d51-b6569f4a917e · inbound

Robust and Label-Efficient Deep Waste Detection cites this paper.

Robust and Label-Efficient Deep Waste Detection Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T16:15:06.970353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:15:06.970353Z digest=sha256:22f3a58afa53a5a42a527966ae30e257b0f92b8d49a962aa58b717344a383640

Observation 2504c725-81df-4bb4-873e-cf83c4670f93 · inbound

VLMs-in-the-Wild: Bridging the Gap Between Academic Benchmarks and Enterprise Reality cites this paper.

VLMs-in-the-Wild: Bridging the Gap Between Academic Benchmarks and Enterprise Reality Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-05T11:11:20.762326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:11:20.762326Z digest=sha256:4d64f6a3cfea03ab1ff7aa80be66f1a3d8ce50c3893837c4bedfccbb350b797b

Observation 0f60d7bd-3257-4cc5-bc51-5fb93d32a14c · inbound

Rate-Distortion Limits for Multimodal Retrieval: Theory, Optimal Codes, and Finite-Sample Guarantees cites this paper.

Rate-Distortion Limits for Multimodal Retrieval: Theory, Optimal Codes, and Finite-Sample Guarantees Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-04T17:18:53.610342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T17:18:53.610342Z digest=sha256:8e8f47e184acdb4104c786d6630f033a6d5f4ab3b0b7475ff5f4e723fe19c684

Observation 8fa62d87-4c65-4139-9306-84d709e30d4f · inbound

The Vision Wormhole: Latent-Space Communication in Heterogeneous Multi-Agent Systems cites this paper.

The Vision Wormhole: Latent-Space Communication in Heterogeneous Multi-Agent Systems Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-02T22:58:07.375228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T22:58:07.375228Z digest=sha256:0792e30e194c072336b1ed145b7040926a2d2606c23e2a2d3247886248c7c175

Observation 79e88e60-3b7d-44f2-a561-87a2df459541 · inbound

WikiCLIP: An Efficient Contrastive Baseline for Open-domain Visual Entity Recognition cites this paper.

WikiCLIP: An Efficient Contrastive Baseline for Open-domain Visual Entity Recognition Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-15T13:15:50.553661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-15T13:11:54.384284Z digest=sha256:de4c09e7beb213d196a101fcb9b73124a31563cb7dd72463ecc7f1be14d26400

Observation 456e6001-3639-4c05-8ea9-b06b04fc89fe · inbound

WikiCLIP: An Efficient Contrastive Baseline for Open-domain Visual Entity Recognition cites this paper.

WikiCLIP: An Efficient Contrastive Baseline for Open-domain Visual Entity Recognition Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-14T23:55:24.006436Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T23:55:24.006436Z digest=sha256:8cd183e90964f63a989ec0c643c7f2da7864e77c8d28d54b8fc911980929dccb

Observation be20fb00-f225-4467-92b9-931a212c9bb9 · inbound

Explicit Logic Channel for Validation and Enhancement of MLLMs on Zero-Shot Tasks cites this paper.

Explicit Logic Channel for Validation and Enhancement of MLLMs on Zero-Shot Tasks Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-21T11:10:02.024909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T11:09:51.816554Z digest=sha256:5278ae1230981df0de2fdc32edc67b22e97c44451e15b7c7d65838c315033d8b

Observation 3d153cec-cc76-4407-83ec-879e1681f651 · inbound

DeCo-DETR: Decoupled Cognition DETR for efficient Open-Vocabulary Object Detection cites this paper.

DeCo-DETR: Decoupled Cognition DETR for efficient Open-Vocabulary Object Detection Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T20:33:16.707134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-13T20:30:16.163179Z digest=sha256:71f470f67963d2d8cb08f7bdfe3fe78a1977b7d9cf9a4d6b37942d991a112248

Observation a5243118-073a-4d3d-9ff6-355111f88bfb · inbound

DeCo-DETR: Decoupled Cognition DETR for efficient Open-Vocabulary Object Detection cites this paper.

DeCo-DETR: Decoupled Cognition DETR for efficient Open-Vocabulary Object Detection Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T21:49:29.323542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-14T21:48:49.892964Z digest=sha256:51cf222d911b25d003a158c226d5529ba7c4604f8fc845254df1748c7e970a56

Observation 2eac7ce6-23f4-4ea8-ba63-a6293d0feb53 · inbound

Latent Anomaly Knowledge Excavation: Unveiling Sparse Sensitive Neurons in Vision-Language Models cites this paper.

Latent Anomaly Knowledge Excavation: Unveiling Sparse Sensitive Neurons in Vision-Language Models Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T06:51:16.216373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-10T17:25:34.942028Z digest=sha256:fbc07ec7402ac2c02ecca8061250f298bf0b5fa0d5c276aa20cf8f6e43499981

Observation 4e82d205-e315-41a4-97c8-1d0208ad0f87 · inbound

Compared to What? Baselines and Metrics for Counterfactual Prompting cites this paper.

Compared to What? Baselines and Metrics for Counterfactual Prompting Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision

Reference 93

Resolution
verified exact
arxiv_id, observed 2026-05-09T19:05:10.548305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-09T19:02:46.991897Z digest=sha256:03e4b0b4a36abec9dde83f486ca81c7fbdc20f32c1df6c7dbe1b68280dbf8014

Observation ef87c4ca-e90d-4d07-a1d1-47f52356d4dc · inbound

Vision Harnessing Agent for Open Ad-hoc Segmentation cites this paper.

Vision Harnessing Agent for Open Ad-hoc Segmentation Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-20T05:53:04.439800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T05:52:40.429412Z digest=sha256:eea63b81d9ca9930673d4bc23d25912dd4ed0981c522c335305dd68c7b287ba8

Observation 44908e21-1cef-439e-b18d-57458df842c3 · inbound

Toward Calibrated, Fair, and accurate Deepfake Detection cites this paper.

Toward Calibrated, Fair, and accurate Deepfake Detection Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision

Reference 266

Resolution
verified exact
arxiv_id, observed 2026-06-28T07:11:45.290744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-06-28T07:05:18.026601Z digest=sha256:9272e5153f7c6cf548c28a54af4b53fefb373415b1b6b9cc7e2bd57ea216f623

Observation 11c0e04f-f341-4fa2-9ced-cdeb5eab031b · inbound

Modeling Complex Behaviors: Multi-Personality Composition and Dynamic Switching in Vision-Language Models cites this paper.

Modeling Complex Behaviors: Multi-Personality Composition and Dynamic Switching in Vision-Language Models Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision

Reference 74

Resolution
verified exact
arxiv_id, observed 2026-07-03T05:47:41.829444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-06-27T13:04:14.886733Z digest=sha256:c3bdf348a9f40441a347a76f036fdb2775de0e1d3668ab16c955c6938e140069

Observation 887c66e1-1378-4117-a8ea-35b4aad639d3 · inbound

Qwen-RobotWorld Technical Report: Unifying Embodied World Modeling through Language-Conditioned Video Generation cites this paper.

Qwen-RobotWorld Technical Report: Unifying Embodied World Modeling through Language-Conditioned Video Generation Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision

Reference 140

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T17:18:43.932557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-06-27T04:19:26.332718Z digest=sha256:e078f0784153ae0a8e81ad22a341ca67b32f55c80e70eef6badf6f52a8c4cf28

Observation daa1c166-07ec-40c9-b3f6-4c193ffb0d88 · inbound

Combating Textual Noise and Redundancy: Entropy-Aware Dense Visual Token Pruning cites this paper.

Combating Textual Noise and Redundancy: Entropy-Aware Dense Visual Token Pruning Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision

Reference 22

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T14:48:32.368383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-03T14:47:35.377391Z digest=sha256:50d8f41ce740df13d1cea3f0f172a0a8c29afc8d6ae1004c26a5d0d53344a3e0

Observation 2a4df5d4-58ff-4181-a3bd-baefb693cf78 · inbound

Qwen-Audio-VAE Technical Report cites this paper.

Qwen-Audio-VAE Technical Report Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision

Reference 155

Resolution
unresolved
no resolver link, observed 2026-07-14T03:31:19.309532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T03:31:19.309532Z digest=sha256:04cd1a254ac7ff665e9500f526e55d258622097343c1722cc85c90143e52558a

Observation d4b3338b-8b46-4a9c-a9ae-6a354c1cdbcc · inbound

Theia: Large-Scale Multimodal Captioning and Automated Validation of the Incidents1M Dataset for Data-Free Distillation cites this paper.

Theia: Large-Scale Multimodal Captioning and Automated Validation of the Incidents1M Dataset for Data-Free Distillation Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-31T12:44:13.340533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T12:44:13.340533Z digest=sha256:34581f1c3f2f7c5b8b957376f12de868e7ea92f9ab8e0d5a94170a4e9616e656

Observation a91ac6c3-9552-4fed-ae18-c46aa3d840f3 · inbound

Same Semantics, Different Paths: Self-Improving Alignment for Vision-Text Compression cites this paper.

Same Semantics, Different Paths: Self-Improving Alignment for Vision-Text Compression Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T15:05:38.990379Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:05:38.990379Z digest=sha256:fce5c9fda1417ba1b8a91915380008ad05298c606cbc4a645aacd3f948f7dfda

Observation e59a47cb-8137-4931-a00c-1b059f3e6e87 · inbound

MASCOT: Model-Aware Submodular Coverage for Composite-Attribute Text-to-Image Retrieval cites this paper.

MASCOT: Model-Aware Submodular Coverage for Composite-Attribute Text-to-Image Retrieval Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-16T00:12:08.581177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:12:08.581177Z digest=sha256:95d8c291dbafa0e0645d272a9c9c643bfb49aa00dfd38108b74ebfbf8809aa7c