Pith. sign in

Paper Citation Record · LEDGER

EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 30 inbound Pith citation observations for arXiv:2402.04252.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2402.04252 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 30 of 30 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 30 of 30 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T00:12:36.489685Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

4
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 95ad16bf-1e74-47a6-9bd8-d43c48855210 · inbound

LR0.FM: Low-Res Benchmark and Improving Robustness for Zero-Shot Classification in Foundation Models cites this paper.

LR0.FM: Low-Res Benchmark and Improving Robustness for Zero-Shot Classification in Foundation Models EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-09T00:12:36.489685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T00:12:36.489685Z digest=sha256:1e693a9c5911793a1e0cd5f5f702aca0a5ff14e390000e8ec0ed268b85671b10

Observation 6bf53bbe-ea34-459b-95a4-8b2c145abedf · inbound

Conformal Predictions for Human Action Recognition with Vision-Language Models cites this paper.

Conformal Predictions for Human Action Recognition with Vision-Language Models EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-08T14:57:12.636325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:57:12.636325Z digest=sha256:cb75e79087078dbf9bb86f03dc1e1d275f31a7f7d4e6321eff48479e2f0f1aa2

Observation 20477537-a323-4c31-956e-b3057dc27188 · inbound

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models cites this paper.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.850157Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.850157Z digest=sha256:3510ebdfd6efc7bb64eb7fdb80272ac07904df727232698151dd252866d5ce49

Observation 761d635e-cc0e-4563-abd0-2d153a7b3767 · inbound

Perception Encoder: The best visual embeddings are not at the output of the network cites this paper.

Perception Encoder: The best visual embeddings are not at the output of the network EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters

Reference 130

Resolution
verified exact
arxiv_id, observed 2026-05-13T22:21:15.804400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T22:21:15.681336Z digest=sha256:086675f7ee2056b7f47549b387468de1a14a59e2882d4f8823be1ed006d963bc

Observation 33163676-9942-43da-96e7-7883f6a5b377 · inbound

Spa-VLM: Stealthy Poisoning Attacks on RAG-based VLM cites this paper.

Spa-VLM: Stealthy Poisoning Attacks on RAG-based VLM EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T13:21:41.212978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:21:41.212978Z digest=sha256:ec522ce9d0329edb25489addaf1c9c77bd8d6295fa02d1c8f539ae25cec6a355

Observation 6141965c-c4e6-409f-9782-80e21e49fed8 · inbound

mRAG: Elucidating the Design Space of Multi-modal Retrieval-Augmented Generation cites this paper.

mRAG: Elucidating the Design Space of Multi-modal Retrieval-Augmented Generation EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T12:42:26.469919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:42:26.469919Z digest=sha256:ab0879570b0e8f019bd4ad63a1798bd2e6e6ac6e2adc42d8f7b22ec0f9c62b32

Observation 4ac52ac6-a3b7-487b-a43c-42578b84fa9e · inbound

Improve Multi-Modal Embedding Learning via Explicit Hard Negative Gradient Amplifying cites this paper.

Improve Multi-Modal Embedding Learning via Explicit Hard Negative Gradient Amplifying EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T13:18:47.669708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:18:47.669708Z digest=sha256:5ee776d9b2bebd40dc3248e3f2edfd1d3c11ce9ed7f5402ce03775fe7db8ad15

Observation b287334b-6846-47e4-9a02-02e8a143c3d6 · inbound

Can Argus Judge Them All? Comparing VLMs Across Domains cites this paper.

Can Argus Judge Them All? Comparing VLMs Across Domains EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:52.199619Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:20:52.199619Z digest=sha256:d3e8af93ea18aec97c924632eadc6e4d65fdaad234e2df8224e69063fa7e2fcf

Observation 4d1c2b4e-a74b-4f70-8c1c-f0b2b2932a61 · inbound

Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames cites this paper.

Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T21:06:26.749785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:06:26.749785Z digest=sha256:3eb5bd338b46974465082a9f58586c98b909f855854eff46f4f844f20db83eb0

Observation 1e092c2f-2333-4d1d-849b-1b18121ed4e0 · inbound

ExpStar: Towards Automatic Commentary Generation for Multi-discipline Scientific Experiments cites this paper.

ExpStar: Towards Automatic Commentary Generation for Multi-discipline Scientific Experiments EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T17:55:45.406624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:55:45.406624Z digest=sha256:8837eac7ab95c460d94305cb03ee4b7b8275df615c300324248d0b1a177693c5

Observation 5006196a-e851-43d4-b838-bb79c4440a6c · inbound

Category-level Text-to-Image Retrieval Improved: Bridging the Domain Gap with Diffusion Models and Vision Encoders cites this paper.

Category-level Text-to-Image Retrieval Improved: Bridging the Domain Gap with Diffusion Models and Vision Encoders EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-05T13:59:44.047838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:59:44.047838Z digest=sha256:39a5cb142d3216a0b2b11cd17ad83aab8ddecc7e08ce773e5cc71cea37e50b17

Observation a5a8c8bd-7223-4a3f-a987-fd10e00d65c9 · inbound

Perceive, Verify and Understand Long Video: Multi-Granular Perception and Active Verification via Interactive Agents cites this paper.

Perceive, Verify and Understand Long Video: Multi-Granular Perception and Active Verification via Interactive Agents EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters

Reference 27

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T12:41:22.579211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-18T12:40:19.544260Z digest=sha256:c2354502d404b7b7bf550433c503c6f09366b6445b8114ebf07b63d0cd4bd05c

Observation c74decde-c081-47ff-baf8-5c04afc10656 · inbound

QKVQA: Question-Focused Filtering for Knowledge-based VQA cites this paper.

QKVQA: Question-Focused Filtering for Knowledge-based VQA EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-16T12:57:54.033143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T12:53:59.668663Z digest=sha256:1d6ecac068fdffb3f060a5c2e980fd4cf51a7b2fc98d4474ea281a9bf29cc886

Observation 24b1a7b5-fabe-48c9-99c0-5ccc3d4e1fe9 · inbound

Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs cites this paper.

Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-03T04:20:57.341823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:20:57.341823Z digest=sha256:224cbe500f1a5e6a920ba4318dfdeb2c7dac2e5fa81ddc5050a7f4d59ee5a852

Observation 210dbf3f-69e6-41c3-ba71-1ed8a84f9b3b · inbound

Explicit Logic Channel for Validation and Enhancement of MLLMs on Zero-Shot Tasks cites this paper.

Explicit Logic Channel for Validation and Enhancement of MLLMs on Zero-Shot Tasks EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-05-21T11:10:02.006151Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-21T11:09:51.816554Z digest=sha256:d9a87b62cc315abcc66ee6eb38bd2ee228b83706b7639faf5299c70ec8143271

Observation a43f5ecc-af12-4c0e-b663-17de2f84e7cd · inbound

Explicit Logic Channel for Validation and Enhancement of MLLMs on Zero-Shot Tasks cites this paper.

Explicit Logic Channel for Validation and Enhancement of MLLMs on Zero-Shot Tasks EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters

Reference 58

Resolution
unresolved
no resolver link, observed 2026-07-14T22:43:59.727378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T22:43:59.727378Z digest=sha256:9af92b7e1dd099c8e13800f966beb2e1335199f8c614cd6e9345541f08e83b13

Observation ab31082e-b61d-4101-a65d-f26a4078c97f · inbound

CoME-VL: Scaling Complementary Multi-Encoder Vision-Language Learning cites this paper.

CoME-VL: Scaling Complementary Multi-Encoder Vision-Language Learning EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-05-13T20:33:14.532847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T20:28:30.864143Z digest=sha256:1c41a3b25dc5779d524395dd6f0b4d24238ddb2bedf2f10652f122bdcd99d7a3

Observation e9dd7262-0e5a-42be-91d3-07b0c2fb289d · inbound

MG$^2$-RAG: Multi-Granularity Graph for Multimodal Retrieval-Augmented Generation cites this paper.

MG$^2$-RAG: Multi-Granularity Graph for Multimodal Retrieval-Augmented Generation EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-05-13T17:13:00.915838Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T17:12:38.618233Z digest=sha256:88fc343c7ce506bea6fa802e2b24991b00e6f108c86f03d6297fdd23863fd40c

Observation 58593182-a8bb-40e1-84b2-f315e94771f7 · inbound

MG$^2$-RAG: Multi-Granularity Graph for Multimodal Retrieval-Augmented Generation cites this paper.

MG$^2$-RAG: Multi-Granularity Graph for Multimodal Retrieval-Augmented Generation EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters

Reference 64

Resolution
unresolved
no resolver link, observed 2026-07-14T19:54:25.294067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T19:54:25.294067Z digest=sha256:333975442beafa08ec842436c17f192041a6cda3e6406b6550bea2fbc0a50d2f

Observation 3f1319da-4fef-4b53-9cca-409ce23938c0 · inbound

Benchmarking Deflection and Hallucination in Large Vision-Language Models cites this paper.

Benchmarking Deflection and Hallucination in Large Vision-Language Models EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters

Reference 4

Resolution
malformed identifier
arxiv_id, observed 2026-05-11T10:31:00.678352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T15:28:26.268341Z digest=sha256:8ae4286b3fd5f82af0f0a9bcd08f3430b360fc68803d2471f133998fc6c006ea

Observation 1e3b25e8-7e0a-4d90-8e9a-b9b75bd515fe · inbound

Chain-of-Models Pre-Training: Rethinking Training Acceleration of Vision Foundation Models cites this paper.

Chain-of-Models Pre-Training: Rethinking Training Acceleration of Vision Foundation Models EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:05:58.682843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T16:16:51.109113Z digest=sha256:fe183576d5a7152ad12cc6fc5bf973c59581b3f1f13d45dc0e3aa066fe6c0015

Observation f31a42a9-3475-4e02-aa22-a9e3b35373a2 · inbound

Exploring High-Order Self-Similarity for Video Understanding cites this paper.

Exploring High-Order Self-Similarity for Video Understanding EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters

Reference 74

Resolution
verified exact
arxiv_id, observed 2026-05-10T00:59:49.461872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T00:59:03.890135Z digest=sha256:ed14836d203a14d3f4dde39575c7a5a3170c78b76ab6a396abaf90b09fb421a4

Observation 5ec127f8-c038-4205-9379-8c004079cb78 · inbound

HiCrew: Hierarchical Reasoning for Long-Form Video Understanding via Question-Aware Multi-Agent Collaboration cites this paper.

HiCrew: Hierarchical Reasoning for Long-Form Video Understanding via Question-Aware Multi-Agent Collaboration EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters

Reference 27

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T14:26:02.222190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-09T21:55:35.699057Z digest=sha256:c230cc19b5f386e1221ed4d0f83b1a501dc1ccfbde060a961efa655ab8de29cd

Observation 1219347e-580a-4b29-802e-2749de453990 · inbound

Sensorimotor World Models: Perception for Action via Inverse Dynamics cites this paper.

Sensorimotor World Models: Perception for Action via Inverse Dynamics EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-07-04T03:09:30.299865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T18:21:51.580903Z digest=sha256:3be8479974d2b2c9aae92c1e58aaa82ee6b60ebd9a2c89ab91b2d335362b015e

Observation f7e01b75-078c-4ca5-8725-73c52626a467 · inbound

ELVA: Exploring Ranking-Driven Universal Multimodal Retrieval cites this paper.

ELVA: Exploring Ranking-Driven Universal Multimodal Retrieval EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-07-04T05:49:36.655493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T15:34:55.062016Z digest=sha256:15230726b3b1fb93d810a168d7c257a0af5fa6b96a2373020deadd1d76b71a74

Observation 7bafcb59-91f1-40f5-bec6-abf94911e215 · inbound

Ground Then Rank: Revisiting Knowledge-Based VQA with Training-Free Entity Identification cites this paper.

Ground Then Rank: Revisiting Knowledge-Based VQA with Training-Free Entity Identification EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-07-04T10:59:46.881141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-26T08:12:14.829556Z digest=sha256:a98f5a926de0f7ef1420fd70c613a90805dc4515ddff0c47df85c1c2f7f46133

Observation 3c2d5fa3-575a-4d5f-a188-1c1710b824b6 · inbound

ReasonCLIP-58M: Visually Grounded Commonsense Reasoning Supervision for CLIP cites this paper.

ReasonCLIP-58M: Visually Grounded Commonsense Reasoning Supervision for CLIP EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters

Reference 80

Resolution
verified exact
arxiv_id, observed 2026-07-04T13:39:50.784830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T05:03:15.044146Z digest=sha256:d9a79dc9c56ff82a587149f0abbf884bd5c2d9ab1f3534b468b4bda556e2c43a

Observation 38118793-1777-41f9-b8f6-6a1acc7ce684 · inbound

Identifying and Resolving Pitfalls of Knowledge-Based VQA Benchmarks: Auditing, Repairing, and Augmenting cites this paper.

Identifying and Resolving Pitfalls of Knowledge-Based VQA Benchmarks: Auditing, Repairing, and Augmenting EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-07-02T19:17:17.749685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-02T19:13:44.015782Z digest=sha256:a3cd58d43d68e7d8ddb7be28319df341d9d59c022f441240bef3d69e461db2f3

Observation 851cac4c-9e90-4435-b87d-30ba397117c2 · inbound

UniHEAR: Unified Heterogeneous-Source Attentive Retrieval for Knowledge-Based Visual Question Answering cites this paper.

UniHEAR: Unified Heterogeneous-Source Attentive Retrieval for Knowledge-Based Visual Question Answering EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T00:31:17.047603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:31:17.047603Z digest=sha256:547c764eee42679c5444453296cc8731070e3a32637992b8e30d55b1650974f3

Observation 353c7071-a8c2-4dbc-9acf-c8a45a9ce76c · inbound

UniHEAR: Unified Heterogeneous-Source Attentive Retrieval for Knowledge-Based Visual Question Answering cites this paper.

UniHEAR: Unified Heterogeneous-Source Attentive Retrieval for Knowledge-Based Visual Question Answering EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T04:19:52.837353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:19:52.837353Z digest=sha256:df3dcf18e4121360c81b487c60d473e898f3a2031b36097c17328f675ca7e74f