Pith. sign in

Paper Citation Record · LEDGER

Meta CLIP 2: A Worldwide Scaling Recipe

As of 20 August 2026, this Paper Citation Record lists 29 of 29 outbound references and 24 inbound Pith citation observations for arXiv:2507.22062.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.22062 v3

Coverage vector

measured 29 of 29 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T12:08:23.842628Z

measured 53 of 53 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 24 of 24 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T14:17:49.408405Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T13:39:50.777855Z

Reference resolution

29 of 29 outbound references displayed

  • verified exact0
  • verified fuzzy5
  • unresolved24
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1d034e13-e2da-40ca-b488-e82fea865f11 · outbound

This paper cites Towards Zero-shot Cross-lingual Image Retrieval.

Meta CLIP 2: A Worldwide Scaling Recipe Towards Zero-shot Cross-lingual Image Retrieval

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T12:08:21.963849Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:08:21.963849Z digest=sha256:0645f583a8d80dab0ec4b13ae27ef57151e0010421f7052729c2eb504867a085

Observation 5c662cf9-bafa-4fd2-9608-654a664a0bd5 · outbound

This paper cites an unresolved cited work.

Meta CLIP 2: A Worldwide Scaling Recipe Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-06T12:08:24.446114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T12:08:23.842628Z digest=sha256:dc8ec0eec2f449c33b541642382cd955ac967154e10a60a44afc1ea4f7bba72f

Observation ff29788c-20c9-4cd0-b73c-f8bff230d0e4 · outbound

This paper cites Learning word vectors for 157 languages.

Meta CLIP 2: A Worldwide Scaling Recipe Learning word vectors for 157 languages

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:08:25.102057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T12:08:22.555512Z digest=sha256:9de981807cee469a6c0892057d530aeff61edb3fcac0c458c2bcbad0a23455f1

Observation 7649df2c-aa6d-447c-96c1-0a43d64893d2 · outbound

This paper cites Graph-RISE: Graph-Regularized Image Semantic Embedding.

Meta CLIP 2: A Worldwide Scaling Recipe Graph-RISE: Graph-Regularized Image Semantic Embedding

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T12:08:22.817078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:08:22.817078Z digest=sha256:5657eba3a31d9a50bb29c44f0dbd566c05dc7237d8b0fb9eccab1f4e1b51a0b7

Observation 75adb5fa-a286-4cdc-b65e-24144d4bfef7 · outbound

This paper cites The open images dataset v4: Unified image classification, object detection, and visual relationship detection at scale.International journal of computer vision , 128(7):1956–1981,.

Meta CLIP 2: A Worldwide Scaling Recipe The open images dataset v4: Unified image classification, object detection, and visual relationship detection at scale.International journal of computer vision , 128(7):1956–1981,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:08:24.913110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T12:08:22.884214Z digest=sha256:bb6aadfadfa9a718a93d15adda9f256ff95bc35e7cda0061763a8ff4fe7ee297

Observation d1d3d36e-405a-4426-bff0-9baa7f1a9f13 · outbound

This paper cites XLM-V: Overcoming the Vocabulary Bottleneck in Multilingual Masked Language Models.

Meta CLIP 2: A Worldwide Scaling Recipe XLM-V: Overcoming the Vocabulary Bottleneck in Multilingual Masked Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T12:08:22.953395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:08:22.953395Z digest=sha256:285aa9cbd258a0f56dc213029675ca2fa4eef6cf8dd88545a66a0d2864c3a363

Observation 026b57e2-414c-42b3-ac8e-6a91ed6ce1f7 · outbound

This paper cites SLIP: Self-supervision meets Language-Image Pre-training.

Meta CLIP 2: A Worldwide Scaling Recipe SLIP: Self-supervision meets Language-Image Pre-training

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T12:08:23.020444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:08:23.020444Z digest=sha256:48a7e0584bb8ef2b6c244541de1de5ca1e531c6c0e1af7b1e83c7e4c89afa502

Observation 31426624-87e8-4e6a-8b42-cb9d16d1b56b · outbound

This paper cites CAPIVARA: Cost-Efficient Approach for Improving Multilingual CLIP Performance on Low-Resource Languages.

Meta CLIP 2: A Worldwide Scaling Recipe CAPIVARA: Cost-Efficient Approach for Improving Multilingual CLIP Performance on Low-Resource Languages

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T12:08:23.119878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:08:23.119878Z digest=sha256:a3d3e1619cf9c10bf4a966c3451944861c53a8d8b140716ad1f4cd4561bdece2

Observation 281372e5-4acb-46cf-b6d6-69c70178464d · outbound

This paper cites LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs.

Meta CLIP 2: A Worldwide Scaling Recipe LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T12:08:23.191317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:08:23.191317Z digest=sha256:18b11634884532289d948920ebae59067a995348f9b4e28b2e9ef9dea29ad6f4

Observation 2ecc4e40-0b38-4074-ad82-498fc78157c8 · outbound

This paper cites No Classification without Representation: Assessing Geodiversity Issues in Open Data Sets for the Developing World.

Meta CLIP 2: A Worldwide Scaling Recipe No Classification without Representation: Assessing Geodiversity Issues in Open Data Sets for the Developing World

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T12:08:23.261268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:08:23.261268Z digest=sha256:10dcce3d921de1b7203efb7978b9cc71cf1c3b7b9217fa0f8c663fe4daa78202

Observation 98d8d6e3-ba55-4421-9c91-ad6d49c8643f · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Meta CLIP 2: A Worldwide Scaling Recipe Gemini: A Family of Highly Capable Multimodal Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T12:08:23.326200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:08:23.326200Z digest=sha256:3fae0da6e42f07080edee3746af42c3c2194180ebd5b63cd8d12461cb4580ef3

Observation fbeabe99-1b29-4ede-8dd0-70a37045b228 · outbound

This paper cites Gemma: Open Models Based on Gemini Research and Technology.

Meta CLIP 2: A Worldwide Scaling Recipe Gemma: Open Models Based on Gemini Research and Technology

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T12:08:23.350759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:08:23.350759Z digest=sha256:9aed6124b6ed58f16af525a8cc8f4ba393eb3e174aa0e388ac1420c46b611539

Observation 8f4f2302-1313-49f0-9a4e-de621667b2b9 · outbound

This paper cites Crossmodal-3600: A massively multilingual multimodal evaluation dataset.

Meta CLIP 2: A Worldwide Scaling Recipe Crossmodal-3600: A massively multilingual multimodal evaluation dataset

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:08:24.745850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T12:08:23.369919Z digest=sha256:5f63e5aaa72810f2776f54cd4b4cedbfdb6a25a767e51f3045d648d829385b8a

Observation f0a212a6-4dde-4bc2-b561-d028cfc6cc4e · outbound

This paper cites Will we run out of data? Limits of LLM scaling based on human-generated data.

Meta CLIP 2: A Worldwide Scaling Recipe Will we run out of data? Limits of LLM scaling based on human-generated data

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T12:08:23.444125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:08:23.444125Z digest=sha256:9ad70ef57e793c70b794647ff97405cedc529f072f1aaf3884bb9d2498dd659b

Observation bfcc85ce-9b18-495c-a5e2-93236e1e3618 · outbound

This paper cites NLLB-CLIP -- train performant multilingual image retrieval model on a budget.

Meta CLIP 2: A Worldwide Scaling Recipe NLLB-CLIP -- train performant multilingual image retrieval model on a budget

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T12:08:23.507484Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:08:23.507484Z digest=sha256:79f932907156a05f9a31a39abeeb5303a7e68ac408a47427ab84e4e955f92e7d

Observation a9d2dc30-1e99-4169-bad4-5885be5b972f · outbound

This paper cites Scaling Pre-training to One Hundred Billion Data for Vision Language Models.

Meta CLIP 2: A Worldwide Scaling Recipe Scaling Pre-training to One Hundred Billion Data for Vision Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T12:08:23.606804Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:08:23.606804Z digest=sha256:6b9f3d4050b97f8fa32cec9fa790ebb66172696bcbb30a329fd0f63665987c8b

Observation 39208fe8-736d-474d-8ff4-cdd9ef151fa0 · outbound

This paper cites Hu Xu, Saining Xie, Xiaoqing Tan, Po-Yao Huang, Russell Howes, Vasu Sharma, Shang-Wen Li, Gargi Ghosh, Luke Zettlemoyer, and Christoph Feichtenhofer.

Meta CLIP 2: A Worldwide Scaling Recipe Hu Xu, Saining Xie, Xiaoqing Tan, Po-Yao Huang, Russell Howes, Vasu Sharma, Shang-Wen Li, Gargi Ghosh, Luke Zettlemoyer, and Christoph Feichtenhofer

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:08:24.592885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T12:08:23.693138Z digest=sha256:655df077250df8908fe2fd686e3bc578501572701cd797d3fe190f094ad208b3

Observation c7322498-8406-416d-bca1-338d1589208f · outbound

This paper cites mT5: A massively multilingual pre-trained text-to-text transformer.

Meta CLIP 2: A Worldwide Scaling Recipe mT5: A massively multilingual pre-trained text-to-text transformer

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T12:08:23.768605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:08:23.768605Z digest=sha256:2c16fc3954d2f3b2d5bcabd07a12b52b4c3902a6165a6b77d4f981957f1df8a9

Observation d444a0ac-efec-4fa8-bf70-ba13f8d9b634 · outbound

This paper cites Scaling Language-Free Visual Representation Learning.

Meta CLIP 2: A Worldwide Scaling Recipe Scaling Language-Free Visual Representation Learning

Reference 2009

Resolution
unresolved
no resolver link, observed 2026-08-06T12:08:22.283340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:08:22.283340Z digest=sha256:3d7d8dadbe6caae325da19bb1912d7ea0f44d1f88dda2a368007ed26676a050b

Observation 1b915d3e-10f0-470c-b0f5-1145895b90ee · outbound

This paper cites Unsupervised Cross-lingual Representation Learning at Scale.

Meta CLIP 2: A Worldwide Scaling Recipe Unsupervised Cross-lingual Representation Learning at Scale

Reference 2015

Resolution
unresolved
no resolver link, observed 2026-08-06T12:08:22.145470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:08:22.145470Z digest=sha256:71a3f988abadc3751f965fca20fb1c791f0dfd048c2c98f7c4ccfa69727957c8

Observation 92b222e6-291a-482a-a18f-034b2357914e · outbound

This paper cites Chameleon: Mixed-Modal Early-Fusion Foundation Models.

Meta CLIP 2: A Worldwide Scaling Recipe Chameleon: Mixed-Modal Early-Fusion Foundation Models

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-06T12:08:23.297514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:08:23.297514Z digest=sha256:ec5cb1a6ca6d220a4a94d2eb04ebb1962894a957d660b946d217187aed29d735

Observation f75b13bc-7463-4314-9967-2e87e14040ed · outbound

This paper cites Distilling the Knowledge in a Neural Network.

Meta CLIP 2: A Worldwide Scaling Recipe Distilling the Knowledge in a Neural Network

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-06T12:08:22.646735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:08:22.646735Z digest=sha256:4f40b9b6d3b2fe444717b1eec9e8ddca25169909684bcc993417ddc60004e9f7

Observation a7c59474-99bf-47ce-b19a-0831ef121dd2 · outbound

This paper cites Imagenet: A large-scale hierarchical image database.

Meta CLIP 2: A Worldwide Scaling Recipe Imagenet: A large-scale hierarchical image database

Reference 2019

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:08:25.248608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T12:08:22.212731Z digest=sha256:e35ec9bf242f3f03960c0328fb3356ddef76702501021524e1300474a36443a2

Observation 9666c542-7268-46b5-be0f-26fd43753bdd · outbound

This paper cites Microsoft COCO Captions: Data Collection and Evaluation Server.

Meta CLIP 2: A Worldwide Scaling Recipe Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-06T12:08:22.071806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:08:22.071806Z digest=sha256:f984767d9e7e1a863eeddaf8df681df8b24cd1fedbd2deca4f37b2b537d9ab98

Observation 86a79199-d5fd-4a11-a37a-b19ab22ad20e · outbound

This paper cites If you use this software, please cite it as below.

Meta CLIP 2: A Worldwide Scaling Recipe If you use this software, please cite it as below

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-06T12:08:22.740432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:08:22.740432Z digest=sha256:fc39bcffe87ec6fbcc56ecc5aa541bda35656532dd38d7c6a6eaa87426586b13

Observation 249240a4-2ef1-4e62-bc20-b991b752b8c4 · outbound

This paper cites SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features.

Meta CLIP 2: A Worldwide Scaling Recipe SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-06T12:08:23.416279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:08:23.416279Z digest=sha256:0a838104eef79530c8574ef0dee379a1fa5df0ad820311379c79df17ca9f2360

Observation 7f6ab0a0-fc9a-4a6a-95b0-28bbcd94b918 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Meta CLIP 2: A Worldwide Scaling Recipe Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-06T12:08:21.997481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:08:21.997481Z digest=sha256:030a0f19e7c614270d82836fc6b1b952db6514b0a506226bfff3f57b6d4f3d6b

Observation ad84ca00-79a6-412f-85d1-f5c3e8ab55c5 · outbound

This paper cites The Llama 3 Herd of Models.

Meta CLIP 2: A Worldwide Scaling Recipe The Llama 3 Herd of Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T12:08:22.487941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:08:22.487941Z digest=sha256:3c41f4b1cee72d16c74250946946af703d570d0f10eb46aa6dc7ee5940a3849a

Observation d8b31b26-ff84-4152-815b-05776a968cd8 · outbound

This paper cites Data Filtering Networks.

Meta CLIP 2: A Worldwide Scaling Recipe Data Filtering Networks

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-06T12:08:22.412069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:08:22.412069Z digest=sha256:df918a9a832816a0e31f3a0087eb205bc5674e55b3f15b9e80a5b07bed1f89e1

Pith citing papers

Observation 5332ff76-ab9f-4b47-a996-dc1aa47a076a · inbound

MetaEmbed: Scaling Multimodal Retrieval at Test-Time with Flexible Late Interaction cites this paper.

MetaEmbed: Scaling Multimodal Retrieval at Test-Time with Flexible Late Interaction Meta CLIP 2: A Worldwide Scaling Recipe

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-18T14:11:27.469427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-18T14:09:22.942238Z digest=sha256:bb1f3638f42c716fd078276e4a36b6d7865084a462398b133ca27534edcbe814

Observation 5ce06e82-7188-48f2-8e50-03db7c0ebcbe · inbound

GRAPE: Let GRPO Supervise Query Rewriting by Ranking for Retrieval cites this paper.

GRAPE: Let GRPO Supervise Query Rewriting by Ranking for Retrieval Meta CLIP 2: A Worldwide Scaling Recipe

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-18T12:21:21.357576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-18T12:18:06.259724Z digest=sha256:ea2b92147d4592306eebdb9fe06c4b703d8fa3d54c6794cd93c4fc02326b12cd

Observation cc627568-7ec6-4a7e-a864-c6d92a5ffca6 · inbound

FG-CLIP 2: A Bilingual Fine-grained Vision-Language Alignment Model cites this paper.

FG-CLIP 2: A Bilingual Fine-grained Vision-Language Alignment Model Meta CLIP 2: A Worldwide Scaling Recipe

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-04T10:18:42.782278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:18:42.782278Z digest=sha256:5912c23d66a7fcd768582c4160d9007411484e20494b10988efea74586dd6f6d

Observation 56e06a15-ef53-492f-a6c2-85132b4de802 · inbound

PowerCLIP: Powerset Alignment for Contrastive Pre-Training cites this paper.

PowerCLIP: Powerset Alignment for Contrastive Pre-Training Meta CLIP 2: A Worldwide Scaling Recipe

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-17T04:59:03.966572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-17T04:57:44.794342Z digest=sha256:9d865e488d45e4838d8acf3a801f3b76a2cfcab73cefdfd841a988986854cd40

Observation 97df5930-a249-47b2-bfd1-82bba7245d8d · inbound

Simplicity Prevails: The Emergence of Generalizable AIGI Detection in Visual Foundation Models cites this paper.

Simplicity Prevails: The Emergence of Generalizable AIGI Detection in Visual Foundation Models Meta CLIP 2: A Worldwide Scaling Recipe

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:40:46.262899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-16T08:40:09.385451Z digest=sha256:8eb434d0009da3134f53173acec058a3624852c323e53f80269eda030077400f

Observation 5cc71124-0f00-41db-97c1-36ec5518baa5 · inbound

Xray-Visual Models: Scaling Vision models on Industry Scale Data cites this paper.

Xray-Visual Models: Scaling Vision models on Industry Scale Data Meta CLIP 2: A Worldwide Scaling Recipe

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-02T22:25:58.596814Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:25:58.596814Z digest=sha256:1a4036a787828eae9ec23f7224785aa88b547bb0f08790c9a6324f6fc434133e

Observation f1a22a00-6cc0-4285-8f90-441f399b94e6 · inbound

Peel neighborhoods cites this paper.

Peel neighborhoods Meta CLIP 2: A Worldwide Scaling Recipe

Reference 22

Resolution
unresolved
no resolver link, observed 2026-07-14T20:08:04.785209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T20:08:04.785209Z digest=sha256:aebb9c50f51fc407d1b889745665ba337635e2bd37d8af605b31054be0ba0562

Observation d49a169b-8bc2-470a-895e-0ce8b4fcf39e · inbound

When Surfaces Lie: Exploiting Wrinkle-Induced Attention Shift to Attack Vision-Language Models cites this paper.

When Surfaces Lie: Exploiting Wrinkle-Induced Attention Shift to Attack Vision-Language Models Meta CLIP 2: A Worldwide Scaling Recipe

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-14T21:19:28.021625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-14T21:19:04.468630Z digest=sha256:ce535408f0383147fa2b1fb439d1c9209ff2b6d955766783118c4b7d85f3b445

Observation 0bda72bd-3ab8-4aa5-8ca3-ab1823d882fc · inbound

HEDGE: Heterogeneous Ensemble for Detection of AI-GEnerated Images in the Wild cites this paper.

HEDGE: Heterogeneous Ensemble for Detection of AI-GEnerated Images in the Wild Meta CLIP 2: A Worldwide Scaling Recipe

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-13T18:58:09.042956Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-13T18:53:14.522561Z digest=sha256:5e82c84aaf2864877304684fa738534d7193c6ff877ce82a459da5c801a507a4

Observation 113b0666-aff2-487b-b8b9-9c5224cb1394 · inbound

LOGER: Local--Global Ensemble for Robust Deepfake Detection in the Wild cites this paper.

LOGER: Local--Global Ensemble for Robust Deepfake Detection in the Wild Meta CLIP 2: A Worldwide Scaling Recipe

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-13T18:38:07.133714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-13T18:37:11.529615Z digest=sha256:f3e0433205ea1f3292672c7912288161e5d4558d707056538a0497de380120d8

Observation c9e66fd2-8d9a-4934-864f-971ae008fe16 · inbound

SearchAD: Large-Scale Rare Image Retrieval Dataset for Autonomous Driving cites this paper.

SearchAD: Large-Scale Rare Image Retrieval Dataset for Autonomous Driving Meta CLIP 2: A Worldwide Scaling Recipe

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-11T07:56:01.620906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-10T16:52:37.723921Z digest=sha256:7801cbe9b9aeeda354f9ea4d9d64ca726741fa073779e5f167e7c2a19386921c

Observation 2842cd12-81a7-4d8a-944c-a8efe8790ae3 · inbound

Boosting Robust AIGI Detection with LoRA-based Pairwise Training cites this paper.

Boosting Robust AIGI Detection with LoRA-based Pairwise Training Meta CLIP 2: A Worldwide Scaling Recipe

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-11T11:26:01.868275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-10T14:56:11.099966Z digest=sha256:361c0a38ff7a0a348e6b2cd83aa62553390d8d5cab551a95f4d9d4ebf349dcaa

Observation a63c8dcf-d310-452d-8cad-dc0de84359df · inbound

Concrete Jungle: Towards Concreteness Paved Contrastive Negative Mining for Compositional Understanding cites this paper.

Concrete Jungle: Towards Concreteness Paved Contrastive Negative Mining for Compositional Understanding Meta CLIP 2: A Worldwide Scaling Recipe

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:36:05.030263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-10T15:24:57.169737Z digest=sha256:364ce6780b6f75856bb1253efcf56eb4d4024fcdbc54e4d5fbfd4a4b2e7707f7

Observation 985bcb05-b485-409e-95e4-6b89a8d86d05 · inbound

Sapiens2 cites this paper.

Sapiens2 Meta CLIP 2: A Worldwide Scaling Recipe

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-11T14:21:06.956239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-09T21:59:43.755956Z digest=sha256:5dddb9e19b083c22f03a37d79a6a7a6d4ead4ee623a2ede3c3c7993eadf64eb6

Observation f69a7296-e811-4d1f-a3a1-280deaa44782 · inbound

FLARE: Full-Modality Long-Video Audiovisual Retrieval Benchmark with User-Simulated Queries cites this paper.

FLARE: Full-Modality Long-Video Audiovisual Retrieval Benchmark with User-Simulated Queries Meta CLIP 2: A Worldwide Scaling Recipe

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:26:25.348924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-12T05:22:14.870351Z digest=sha256:498f22a180114f5eb7d9f2e1a34d7305974983ac2643074ba2f8da49467b01c9

Observation 3eaeb297-59c8-4cec-ab70-77fe8fb0f498 · inbound

CRAFT: Clinical Reward-Aligned Finetuning for Medical Image Synthesis cites this paper.

CRAFT: Clinical Reward-Aligned Finetuning for Medical Image Synthesis Meta CLIP 2: A Worldwide Scaling Recipe

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-14T21:18:00.153547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-14T21:10:36.000486Z digest=sha256:b3bbb9574faa4308caa15710a0fe92414452964aa11eeafad627d279be54b3e8

Observation 2d0bd3ec-2eeb-4629-abfb-16f5a8519ac1 · inbound

Beyond Symmetric Alignment: Spectral Diagnostics of Modality Imbalance in Vision-Language Models in the Medical Domain cites this paper.

Beyond Symmetric Alignment: Spectral Diagnostics of Modality Imbalance in Vision-Language Models in the Medical Domain Meta CLIP 2: A Worldwide Scaling Recipe

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-07-02T07:46:46.323860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-28T06:40:00.643198Z digest=sha256:82aae21877949146e980fde4c2b461cb515c4b8e5b6f01a71b91bcc305d61057

Observation e3f4c34d-e547-49c8-bfe1-d21041147b67 · inbound

ReasonCLIP-58M: Visually Grounded Commonsense Reasoning Supervision for CLIP cites this paper.

ReasonCLIP-58M: Visually Grounded Commonsense Reasoning Supervision for CLIP Meta CLIP 2: A Worldwide Scaling Recipe

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T13:39:50.779287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-26T05:03:15.044146Z digest=sha256:cbc177df78ed3589591c4b46ed039decd95fbaa6021b66ab1fee5f3b06f26bb9

Observation 9b619a9b-4c31-4cf0-bcb1-e3e75745f5fa · inbound

AdaBoosting Text Prompts for Vision-Language Models cites this paper.

AdaBoosting Text Prompts for Vision-Language Models Meta CLIP 2: A Worldwide Scaling Recipe

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T16:27:08.304554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-02T16:24:03.231930Z digest=sha256:646a5be56355051def59aa31b42140246b361ff10ed50b32a7144e99e2ac91b9

Observation 85cbf92b-cce4-4563-9cfe-942b987fdf07 · inbound

Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model cites this paper.

Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model Meta CLIP 2: A Worldwide Scaling Recipe

Reference 106

Resolution
unresolved
no resolver link, observed 2026-07-31T06:20:14.029211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:20:14.029211Z digest=sha256:dd9cd68a5144d1f6f728fdb3d53fe667da3d088bdd2c0e4109c2abefaf4a745d

Observation 6d0e6292-82af-4548-95f5-d21db5e3b9d6 · inbound

Fine-Grained Food Image Understanding via Target-Aware Data Alignment cites this paper.

Fine-Grained Food Image Understanding via Target-Aware Data Alignment Meta CLIP 2: A Worldwide Scaling Recipe

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-01T01:32:32.988872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T01:32:32.988872Z digest=sha256:ad30215fb2ac5f0dceda0f94524c3273387633c0d7a5ca4947fe0654110c7783

Observation 5c7b89bf-88f1-4b03-9f7e-bfa49d0f73af · inbound

Enhancing VLM Reward Models Through Structure-Aware Fine-Tuning cites this paper.

Enhancing VLM Reward Models Through Structure-Aware Fine-Tuning Meta CLIP 2: A Worldwide Scaling Recipe

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-05T10:39:50.676419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:39:50.676419Z digest=sha256:22bd30a35a5eb8abf0f81c0116ad367ffdca0bbcd34e01ef397e9ab3412ff74a

Observation c5c283fd-8101-4f98-96cc-19bf708b8a8d · inbound

On the Limitations of Cross-Lingual Consistency in Multilingual Text-to-image Generation cites this paper.

On the Limitations of Cross-Lingual Consistency in Multilingual Text-to-image Generation Meta CLIP 2: A Worldwide Scaling Recipe

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T12:27:44.166241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:27:44.166241Z digest=sha256:13709842029232a0a8808d20550d36d6590f6a3c071eabf0cb2eb0aba00e2eb2

Observation f5a857dc-c433-4b94-ad53-eb73b7d7f29c · inbound

Gaze Target Estimation Anywhere with Concepts cites this paper.

Gaze Target Estimation Anywhere with Concepts Meta CLIP 2: A Worldwide Scaling Recipe

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T14:17:49.408405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:17:49.408405Z digest=sha256:5b2ac545f646e3fa4b38edf5ce5e6b17b3a363b331881e26d79ff26d0c4a66cd