Pith. sign in

Paper Citation Record · LEDGER

EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 30 inbound Pith citation observations for arXiv:2402.04252.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2402.04252 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 30 of 30 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 30 of 30 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T00:12:36.489685Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

4
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 95ad16bf-1e74-47a6-9bd8-d43c48855210 · inbound

LR0.FM: Low-Res Benchmark and Improving Robustness for Zero-Shot Classification in Foundation Models cites this paper.

LR0.FM: Low-Res Benchmark and Improving Robustness for Zero-Shot Classification in Foundation Models EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-09T00:12:36.489685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T00:12:36.489685Z digest=sha256:15a896a7bbbf4006998268509e98597c6e4d2f07a002ed2c389919c534f594c2

Observation 6bf53bbe-ea34-459b-95a4-8b2c145abedf · inbound

Conformal Predictions for Human Action Recognition with Vision-Language Models cites this paper.

Conformal Predictions for Human Action Recognition with Vision-Language Models EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-08T14:57:12.636325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:57:12.636325Z digest=sha256:b78631256115063b7921fcdd2a85196b4ce0bd8d52d2615ea5c7d7cd87996668

Observation 20477537-a323-4c31-956e-b3057dc27188 · inbound

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models cites this paper.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.850157Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.850157Z digest=sha256:3f4b9af57408c8a8fbc17b720cfa2f433b903656b7df0e5c3c227851eaf774b7

Observation 761d635e-cc0e-4563-abd0-2d153a7b3767 · inbound

Perception Encoder: The best visual embeddings are not at the output of the network cites this paper.

Perception Encoder: The best visual embeddings are not at the output of the network EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters

Reference 130

Resolution
verified exact
arxiv_id, observed 2026-05-13T22:21:15.804400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T22:21:15.681336Z digest=sha256:76e06fd2f17abb74e91009ed0b5c9c8901b97b4841cfe5e1d0a1c287078876d9

Observation 33163676-9942-43da-96e7-7883f6a5b377 · inbound

Spa-VLM: Stealthy Poisoning Attacks on RAG-based VLM cites this paper.

Spa-VLM: Stealthy Poisoning Attacks on RAG-based VLM EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T13:21:41.212978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:21:41.212978Z digest=sha256:2982fb67f7d577456e33b4f10989f35ba8eb5d0e90d68d10ddca283678994ba0

Observation 6141965c-c4e6-409f-9782-80e21e49fed8 · inbound

mRAG: Elucidating the Design Space of Multi-modal Retrieval-Augmented Generation cites this paper.

mRAG: Elucidating the Design Space of Multi-modal Retrieval-Augmented Generation EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T12:42:26.469919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:42:26.469919Z digest=sha256:273baffd7fa62d772508bf018d167a5edc81a65ff47c9842aedfc4947246cb7d

Observation 4ac52ac6-a3b7-487b-a43c-42578b84fa9e · inbound

Improve Multi-Modal Embedding Learning via Explicit Hard Negative Gradient Amplifying cites this paper.

Improve Multi-Modal Embedding Learning via Explicit Hard Negative Gradient Amplifying EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T13:18:47.669708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:18:47.669708Z digest=sha256:4062585af9b6418ef755bfb7860b0374e71ac038083ebce67f1ede4e10dccbdc

Observation b287334b-6846-47e4-9a02-02e8a143c3d6 · inbound

Can Argus Judge Them All? Comparing VLMs Across Domains cites this paper.

Can Argus Judge Them All? Comparing VLMs Across Domains EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:52.199619Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:20:52.199619Z digest=sha256:222dc7d7a30c9e6de2b8fad9b920998aeb9a534e0517a1957e0e9f4406e57ed7

Observation 4d1c2b4e-a74b-4f70-8c1c-f0b2b2932a61 · inbound

Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames cites this paper.

Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T21:06:26.749785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:06:26.749785Z digest=sha256:1cde67b12352ec08fdcd3b1a777d5cc9a35a1cc7862f43aa1ca50ba07802e6e4

Observation 1e092c2f-2333-4d1d-849b-1b18121ed4e0 · inbound

ExpStar: Towards Automatic Commentary Generation for Multi-discipline Scientific Experiments cites this paper.

ExpStar: Towards Automatic Commentary Generation for Multi-discipline Scientific Experiments EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T17:55:45.406624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:55:45.406624Z digest=sha256:79bac9726ff3cd99ed1f371fa3a235fdf38f2dd3280e3c787d12c7655bef9930

Observation 5006196a-e851-43d4-b838-bb79c4440a6c · inbound

Category-level Text-to-Image Retrieval Improved: Bridging the Domain Gap with Diffusion Models and Vision Encoders cites this paper.

Category-level Text-to-Image Retrieval Improved: Bridging the Domain Gap with Diffusion Models and Vision Encoders EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-05T13:59:44.047838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:59:44.047838Z digest=sha256:867425b6540b2c89ae03417080cb3a0fc9be8f836f1cceec0a643cb6d9977db9

Observation a5a8c8bd-7223-4a3f-a987-fd10e00d65c9 · inbound

Perceive, Verify and Understand Long Video: Multi-Granular Perception and Active Verification via Interactive Agents cites this paper.

Perceive, Verify and Understand Long Video: Multi-Granular Perception and Active Verification via Interactive Agents EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters

Reference 27

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T12:41:22.579211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-18T12:40:19.544260Z digest=sha256:d29ed5c3c0be2f0049e5bdcda14a08464b95f367f091841a15c6d5160a9ee721

Observation c74decde-c081-47ff-baf8-5c04afc10656 · inbound

QKVQA: Question-Focused Filtering for Knowledge-based VQA cites this paper.

QKVQA: Question-Focused Filtering for Knowledge-based VQA EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-16T12:57:54.033143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T12:53:59.668663Z digest=sha256:b2cbcc020bd96f15a09f46ef14630dfcdb16219a3eed27c91ceb5fdd5be350e0

Observation 24b1a7b5-fabe-48c9-99c0-5ccc3d4e1fe9 · inbound

Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs cites this paper.

Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-03T04:20:57.341823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:20:57.341823Z digest=sha256:cd255ac606239087cf62ea3aeecd7d6e8ce1277f9f16fc0f0626273905d6832e

Observation 210dbf3f-69e6-41c3-ba71-1ed8a84f9b3b · inbound

Explicit Logic Channel for Validation and Enhancement of MLLMs on Zero-Shot Tasks cites this paper.

Explicit Logic Channel for Validation and Enhancement of MLLMs on Zero-Shot Tasks EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-05-21T11:10:02.006151Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-21T11:09:51.816554Z digest=sha256:b7c5891d469a2cd6b8d76016d0b79420bf6fb196d6bc1b7d74502fc7db15104a

Observation a43f5ecc-af12-4c0e-b663-17de2f84e7cd · inbound

Explicit Logic Channel for Validation and Enhancement of MLLMs on Zero-Shot Tasks cites this paper.

Explicit Logic Channel for Validation and Enhancement of MLLMs on Zero-Shot Tasks EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters

Reference 58

Resolution
unresolved
no resolver link, observed 2026-07-14T22:43:59.727378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T22:43:59.727378Z digest=sha256:34b64e74aea73be71febc085a5aebeb6a416ea8e2d2969d1d7a376f63d97ac58

Observation ab31082e-b61d-4101-a65d-f26a4078c97f · inbound

CoME-VL: Scaling Complementary Multi-Encoder Vision-Language Learning cites this paper.

CoME-VL: Scaling Complementary Multi-Encoder Vision-Language Learning EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-05-13T20:33:14.532847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T20:28:30.864143Z digest=sha256:f7bfcd741edf57a442d3635f47008ea4dc585b01db2eaa0b748d77c343ae9f79

Observation e9dd7262-0e5a-42be-91d3-07b0c2fb289d · inbound

MG$^2$-RAG: Multi-Granularity Graph for Multimodal Retrieval-Augmented Generation cites this paper.

MG$^2$-RAG: Multi-Granularity Graph for Multimodal Retrieval-Augmented Generation EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-05-13T17:13:00.915838Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T17:12:38.618233Z digest=sha256:ab5b56fe8c9f1d1d7866fdca0e593335da7129eb3ef3d07f52a1d8bd181d72b0

Observation 58593182-a8bb-40e1-84b2-f315e94771f7 · inbound

MG$^2$-RAG: Multi-Granularity Graph for Multimodal Retrieval-Augmented Generation cites this paper.

MG$^2$-RAG: Multi-Granularity Graph for Multimodal Retrieval-Augmented Generation EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters

Reference 64

Resolution
unresolved
no resolver link, observed 2026-07-14T19:54:25.294067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T19:54:25.294067Z digest=sha256:8f3782138e794b0b2bb77676c1f3baeb214c0729e9f83ab933ebf162666fa1db

Observation 3f1319da-4fef-4b53-9cca-409ce23938c0 · inbound

Benchmarking Deflection and Hallucination in Large Vision-Language Models cites this paper.

Benchmarking Deflection and Hallucination in Large Vision-Language Models EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters

Reference 4

Resolution
malformed identifier
arxiv_id, observed 2026-05-11T10:31:00.678352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T15:28:26.268341Z digest=sha256:18d4ec3ace65114e44579c0f29c37e1c4c7f4378cf407bb7e069deb4fb4539a6

Observation 1e3b25e8-7e0a-4d90-8e9a-b9b75bd515fe · inbound

Chain-of-Models Pre-Training: Rethinking Training Acceleration of Vision Foundation Models cites this paper.

Chain-of-Models Pre-Training: Rethinking Training Acceleration of Vision Foundation Models EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:05:58.682843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T16:16:51.109113Z digest=sha256:3d85dc831e22c39e9dca086d909d60d0be1f20c3953acf3d77ed86bbeac21440

Observation f31a42a9-3475-4e02-aa22-a9e3b35373a2 · inbound

Exploring High-Order Self-Similarity for Video Understanding cites this paper.

Exploring High-Order Self-Similarity for Video Understanding EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters

Reference 74

Resolution
verified exact
arxiv_id, observed 2026-05-10T00:59:49.461872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T00:59:03.890135Z digest=sha256:ec948cb46015d891e88249efe0a6b7e1204de1234857e8bd0b3a6a3cdff2b3cb

Observation 5ec127f8-c038-4205-9379-8c004079cb78 · inbound

HiCrew: Hierarchical Reasoning for Long-Form Video Understanding via Question-Aware Multi-Agent Collaboration cites this paper.

HiCrew: Hierarchical Reasoning for Long-Form Video Understanding via Question-Aware Multi-Agent Collaboration EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters

Reference 27

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T14:26:02.222190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-09T21:55:35.699057Z digest=sha256:6a239d9fea2bda7ff070f0f98f1766535fed42a8895d018ab081128e2a057824

Observation 1219347e-580a-4b29-802e-2749de453990 · inbound

Sensorimotor World Models: Perception for Action via Inverse Dynamics cites this paper.

Sensorimotor World Models: Perception for Action via Inverse Dynamics EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-07-04T03:09:30.299865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T18:21:51.580903Z digest=sha256:5099bc2728b914a174e6737ce47eddd3c222b3ffa500e97a54b4dc57e4ca3280

Observation f7e01b75-078c-4ca5-8725-73c52626a467 · inbound

ELVA: Exploring Ranking-Driven Universal Multimodal Retrieval cites this paper.

ELVA: Exploring Ranking-Driven Universal Multimodal Retrieval EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-07-04T05:49:36.655493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T15:34:55.062016Z digest=sha256:4e4193914b22ee0d5b36a484d788d70179082628b7823a96eacceff999bba62c

Observation 7bafcb59-91f1-40f5-bec6-abf94911e215 · inbound

Ground Then Rank: Revisiting Knowledge-Based VQA with Training-Free Entity Identification cites this paper.

Ground Then Rank: Revisiting Knowledge-Based VQA with Training-Free Entity Identification EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-07-04T10:59:46.881141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-26T08:12:14.829556Z digest=sha256:c88f25bd7a7e111322bd9cbeb84cc84fdb3717dc759a8057354aecbf303d7819

Observation 3c2d5fa3-575a-4d5f-a188-1c1710b824b6 · inbound

ReasonCLIP-58M: Visually Grounded Commonsense Reasoning Supervision for CLIP cites this paper.

ReasonCLIP-58M: Visually Grounded Commonsense Reasoning Supervision for CLIP EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters

Reference 80

Resolution
verified exact
arxiv_id, observed 2026-07-04T13:39:50.784830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T05:03:15.044146Z digest=sha256:a7a032edae18f5da8cc52c94f15cce2c062cc4c1f50cefc79be8f0f4fbfc45cc

Observation 38118793-1777-41f9-b8f6-6a1acc7ce684 · inbound

Identifying and Resolving Pitfalls of Knowledge-Based VQA Benchmarks: Auditing, Repairing, and Augmenting cites this paper.

Identifying and Resolving Pitfalls of Knowledge-Based VQA Benchmarks: Auditing, Repairing, and Augmenting EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-07-02T19:17:17.749685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-02T19:13:44.015782Z digest=sha256:bf59885188e3b4ce3953517a55564956c917399d0d3a176783f31510754b755e

Observation 851cac4c-9e90-4435-b87d-30ba397117c2 · inbound

UniHEAR: Unified Heterogeneous-Source Attentive Retrieval for Knowledge-Based Visual Question Answering cites this paper.

UniHEAR: Unified Heterogeneous-Source Attentive Retrieval for Knowledge-Based Visual Question Answering EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T00:31:17.047603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:31:17.047603Z digest=sha256:f05b0a33a334a8c1df684e1220bd8113c4e178bada48103b151661913a0fb827

Observation 353c7071-a8c2-4dbc-9acf-c8a45a9ce76c · inbound

UniHEAR: Unified Heterogeneous-Source Attentive Retrieval for Knowledge-Based Visual Question Answering cites this paper.

UniHEAR: Unified Heterogeneous-Source Attentive Retrieval for Knowledge-Based Visual Question Answering EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T04:19:52.837353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:19:52.837353Z digest=sha256:5b892c9d47446c7b4f3552e72b1de7103af0db6c3c26a72816da762d218d7ac0