Pith. sign in

Paper Citation Record · LEDGER

From CLIP to DINO: Visual Encoders Shout in Multi-modal Large Language Models

As of 18 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 43 inbound Pith citation observations for arXiv:2310.08825.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2310.08825 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 43 of 43 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 43 of 43 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T15:10:13.086299Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T16:49:58.025162Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation c5cea414-c347-43ed-8d89-640ad01738a2 · inbound

Advancing Fine-Grained Visual Understanding with Multi-Scale Alignment in Multi-Modal Models cites this paper.

Advancing Fine-Grained Visual Understanding with Multi-Scale Alignment in Multi-Modal Models From CLIP to DINO: Visual Encoders Shout in Multi-modal Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T20:27:19.970450Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:27:19.970450Z digest=sha256:fba77cbfaf0954b9b697265ff36120ee79788fea43e60c54239099c561dd3cfc

Observation 87f820bf-13af-4c63-9f8a-c55ea640ef73 · inbound

Libra: Leveraging Temporal Images for Biomedical Radiology Analysis cites this paper.

Libra: Leveraging Temporal Images for Biomedical Radiology Analysis From CLIP to DINO: Visual Encoders Shout in Multi-modal Large Language Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-12T10:20:44.821297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T10:20:44.821297Z digest=sha256:554b96a75a75627710133a6c0975ec12c2799050cccbd95bc03a9c0e282319a5

Observation d7a15a1a-c8b9-4686-adfe-973c94009a05 · inbound

VidHalluc: Evaluating Temporal Hallucinations in Multimodal Large Language Models for Video Understanding cites this paper.

VidHalluc: Evaluating Temporal Hallucinations in Multimodal Large Language Models for Video Understanding From CLIP to DINO: Visual Encoders Shout in Multi-modal Large Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T22:14:09.681095Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:14:09.681095Z digest=sha256:8847e68860e26b375e6eb1f910289df9b3cea54b03facfd33fe4743f390fe2b8

Observation b45d62d5-a49d-40b9-925b-a6e2f82afa10 · inbound

CPath-Omni: A Unified Multimodal Foundation Model for Patch and Whole Slide Image Analysis in Computational Pathology cites this paper.

CPath-Omni: A Unified Multimodal Foundation Model for Patch and Whole Slide Image Analysis in Computational Pathology From CLIP to DINO: Visual Encoders Shout in Multi-modal Large Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T14:21:46.527734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:21:46.527734Z digest=sha256:f0f33a78db84b7a7d7500832daa30f7d93afcf5459d9ab7d451f295f1eeedf49

Observation f191937e-84f9-4da9-bf32-adca4d8e3361 · inbound

ComprehendEdit: A Comprehensive Dataset and Evaluation Framework for Multimodal Knowledge Editing cites this paper.

ComprehendEdit: A Comprehensive Dataset and Evaluation Framework for Multimodal Knowledge Editing From CLIP to DINO: Visual Encoders Shout in Multi-modal Large Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T13:47:10.243345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:47:10.243345Z digest=sha256:1d8022c4821570a91f5890e860c25095b0b2a2c3baa8a2e5423a6ee57497c0a9

Observation 4bd58b9c-39b1-48e2-ad57-c9350f2b7e0d · inbound

Interpretable Face Anti-Spoofing: Enhancing Generalization with Multimodal Large Language Models cites this paper.

Interpretable Face Anti-Spoofing: Enhancing Generalization with Multimodal Large Language Models From CLIP to DINO: Visual Encoders Shout in Multi-modal Large Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T22:25:14.324383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:25:14.324383Z digest=sha256:681be4a60b9140e4d54716d286f745a1826e311e4af0b57c27a7c068a3ce5115

Observation c7b5b60e-23e2-47cf-953a-30ef807b3ed0 · inbound

Diffusion Instruction Tuning cites this paper.

Diffusion Instruction Tuning From CLIP to DINO: Visual Encoders Shout in Multi-modal Large Language Models

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-09T11:20:02.875384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:20:02.875384Z digest=sha256:a9bf8f5815a5284d10234eabb289f8551bf3c25c0ba41d2a9a5b741c7674d75d

Observation 1b2c9241-8d5b-4c41-983b-dcd4728b5c64 · inbound

Toward Generalizable Forgery Detection and Reasoning cites this paper.

Toward Generalizable Forgery Detection and Reasoning From CLIP to DINO: Visual Encoders Shout in Multi-modal Large Language Models

Reference 77

Resolution
verified exact
arxiv_id, observed 2026-05-22T22:27:12.473517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-22T22:26:49.062978Z digest=sha256:4e7aecd1cf90f40c7dfe249f8106975e0920f4b6367adc0330466d742233b7c3

Observation 4372cebe-1cb5-48a4-8acd-b28cb473a856 · inbound

RA-RRG: Multimodal Retrieval-Augmented Radiology Report Generation with Key Phrase Extraction cites this paper.

RA-RRG: Multimodal Retrieval-Augmented Radiology Report Generation with Key Phrase Extraction From CLIP to DINO: Visual Encoders Shout in Multi-modal Large Language Models

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-22T21:12:08.638336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-22T21:09:07.088042Z digest=sha256:190d6cd6c3e9364cf3a3d1b718d6afd6595a2104fa1d3b7ae1b81f352d9b514b

Observation dbba5313-2aad-4768-8d8a-e44aa9ffa657 · inbound

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models cites this paper.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models From CLIP to DINO: Visual Encoders Shout in Multi-modal Large Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:13.676257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:13.676257Z digest=sha256:2450b0d7bf08e7b17de9cb384e73fe142a3307426a3bfa9988a6885c45a8d6f5

Observation ea4cd2ef-2034-4a1b-8b7a-f9b908f0b8ea · inbound

Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor cites this paper.

Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor From CLIP to DINO: Visual Encoders Shout in Multi-modal Large Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T18:52:31.897432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:52:31.897432Z digest=sha256:15b5f54e760195d15e56964730a20f17dfe0d856e00066311c14b4f23e2a0a61

Observation 31d9f30d-3254-4f1c-8878-c08df639758a · inbound

METEOR: Multi-Encoder Collaborative Token Pruning for Efficient Vision Language Models cites this paper.

METEOR: Multi-Encoder Collaborative Token Pruning for Efficient Vision Language Models From CLIP to DINO: Visual Encoders Shout in Multi-modal Large Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T13:18:01.907280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:18:01.907280Z digest=sha256:5aafecc268ea3c8a2041e4b9415f55bf28bacb3c2289ddcbb634f266f32886f8

Observation 6d5705cf-9e95-4c13-baa3-b7bdb6ecc9bd · inbound

CompressKV: Semantic Retrieval Heads Know What Tokens are Not Important Before Generation cites this paper.

CompressKV: Semantic Retrieval Heads Know What Tokens are Not Important Before Generation From CLIP to DINO: Visual Encoders Shout in Multi-modal Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T05:00:04.849995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:00:04.849995Z digest=sha256:0564706c4351ad3d2a0b001471f422e8e1778ba74c81fbbf98bc3458a171b18d

Observation 7de859b7-0815-4f8a-b1e1-53695a0ea7af · inbound

HEAL: A Hypothesis-Based Preference-Aware Analysis Framework cites this paper.

HEAL: A Hypothesis-Based Preference-Aware Analysis Framework From CLIP to DINO: Visual Encoders Shout in Multi-modal Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T15:25:35.340576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:25:35.340576Z digest=sha256:55db810ef589437a997ad9219b638581ebe07770837dcce06cf70b49a3bb88d0

Observation 6c61d66c-4586-4d15-846f-be83e0498925 · inbound

Decoding Memories: An Efficient Pipeline for Self-Consistency Hallucination Detection cites this paper.

Decoding Memories: An Efficient Pipeline for Self-Consistency Hallucination Detection From CLIP to DINO: Visual Encoders Shout in Multi-modal Large Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T14:33:18.602691Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:33:18.602691Z digest=sha256:5d00494923fa5f26473b8a5b0e171df42301ae5bdb0440541edacb62099b3b56

Observation a92a6dea-f979-4319-b35a-d63074e2910a · inbound

TMUAD: Enhancing Logical Capabilities in Unified Anomaly Detection Models with a Text Memory Bank cites this paper.

TMUAD: Enhancing Logical Capabilities in Unified Anomaly Detection Models with a Text Memory Bank From CLIP to DINO: Visual Encoders Shout in Multi-modal Large Language Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-05T14:01:40.129707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:01:40.129707Z digest=sha256:f8483eec02ac8c4696cd746dd78e606f551750349c8b477a23716dacccee8b58

Observation a54931fd-ea48-473d-a5de-b2c9e84bb2d0 · inbound

Testing for LLM response differences: the case of a composite null consisting of semantically irrelevant query perturbations cites this paper.

Testing for LLM response differences: the case of a composite null consisting of semantically irrelevant query perturbations From CLIP to DINO: Visual Encoders Shout in Multi-modal Large Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T17:26:38.328095Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T17:26:38.328095Z digest=sha256:d37f811a2b1a2537402b74beecd36bc3808a610fcbcd6a1b8f0cc2caf008b59b

Observation 41b63a2f-a473-4cd1-a457-f6f7c807a93b · inbound

NP-LoRA: Null Space Projection for Subject-Style LoRA Fusion cites this paper.

NP-LoRA: Null Space Projection for Subject-Style LoRA Fusion From CLIP to DINO: Visual Encoders Shout in Multi-modal Large Language Models

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-25T08:00:29.744034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-25T07:58:25.211030Z digest=sha256:6f89af1cc5d4465b745f3807e38abe053b4ed210ecdffe170acf47404b59c82e

Observation da0fb4f2-cd52-41d2-a5f9-f584920f3438 · inbound

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs cites this paper.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs From CLIP to DINO: Visual Encoders Shout in Multi-modal Large Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:03.179935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:03.179935Z digest=sha256:0c3b24978c53bc125afb8889cb9fbc4c599271dd8de8c8c24c90c0451077af51

Observation c2906415-2ba2-4c97-b381-bcdb3c75e281 · inbound

Neuro-Symbolic Control with Large Language Models for Language-Guided Spatial Tasks cites this paper.

Neuro-Symbolic Control with Large Language Models for Language-Guided Spatial Tasks From CLIP to DINO: Visual Encoders Shout in Multi-modal Large Language Models

Reference 44

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T21:21:17.260679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-16T21:19:37.970460Z digest=sha256:5a4f394e0897a786c884fcdd90624539c26a8b47774959e80e56a6afc64d4833

Observation 008c6eab-522e-4181-8713-264b2bb71c45 · inbound

Who Endorsed It? Measuring Authority Bias Across Expertise Levels in Language Models cites this paper.

Who Endorsed It? Measuring Authority Bias Across Expertise Levels in Language Models From CLIP to DINO: Visual Encoders Shout in Multi-modal Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-03T09:37:07.017396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T09:37:07.017396Z digest=sha256:47d347d06809c9f7dc684da256aa4d8dee20eb9b092aa8a3e04112cb613ccd75

Observation f10a2e3c-0f71-4694-8865-08eabaf7a67d · inbound

Harmonia: Algorithm-Hardware Co-Design for Memory- and Compute-Efficient BFP-based LLM Inference cites this paper.

Harmonia: Algorithm-Hardware Co-Design for Memory- and Compute-Efficient BFP-based LLM Inference From CLIP to DINO: Visual Encoders Shout in Multi-modal Large Language Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-03T04:37:27.791489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:37:27.791489Z digest=sha256:80fa230188113015de61ece8bf7cb702c372fbc5f5940f0fde46d20712fa1d36

Observation 7b8810df-b7bb-4a6a-85ec-1bd91ca6ba43 · inbound

Less is Enough: Synthesizing Diverse Data in LLM Feature Space with Sparse Autoencoders cites this paper.

Less is Enough: Synthesizing Diverse Data in LLM Feature Space with Sparse Autoencoders From CLIP to DINO: Visual Encoders Shout in Multi-modal Large Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-03T01:17:11.852907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:17:11.852907Z digest=sha256:53aa57ee0058b0eebc65bcdfd56e0c8d4aa1836441187e1f36d0fccfbe0eee48

Observation bba09ca1-9f4f-42cc-8fde-16e58a55610e · inbound

InCoM: Intent-Driven Perception and Structured Coordination for Mobile Manipulation cites this paper.

InCoM: Intent-Driven Perception and Structured Coordination for Mobile Manipulation From CLIP to DINO: Visual Encoders Shout in Multi-modal Large Language Models

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-15T19:10:15.922063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-15T19:07:46.733450Z digest=sha256:220ba45f3cfcc9ccb97915b1b36a1dcf6760d7df11ddf00623dcafdee13b2a7e

Observation 14df1184-58e6-4ca9-a7e3-8590e06f4864 · inbound

InCoM: Intent-Driven Perception and Structured Coordination for Mobile Manipulation cites this paper.

InCoM: Intent-Driven Perception and Structured Coordination for Mobile Manipulation From CLIP to DINO: Visual Encoders Shout in Multi-modal Large Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-02T20:35:05.020062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:35:05.020062Z digest=sha256:5b38acdf408e4551ddbf303502813bcb98361df89ef4a450a4a69ee2b7498a15

Observation cfd77212-2fac-4242-bb44-7c4465f2cec7 · inbound

SpatialStack: Layered Geometry-Language Fusion for 3D VLM Spatial Reasoning cites this paper.

SpatialStack: Layered Geometry-Language Fusion for 3D VLM Spatial Reasoning From CLIP to DINO: Visual Encoders Shout in Multi-modal Large Language Models

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-14T22:08:04.602770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-14T22:04:18.591594Z digest=sha256:a7d441415fd2f3832a0ce5b74d21c2e2f1bdada9d64ed2d458d85cc580cea717

Observation 6ab8a71c-97e8-40cf-87fd-de9bf6a49861 · inbound

CoME-VL: Scaling Complementary Multi-Encoder Vision-Language Learning cites this paper.

CoME-VL: Scaling Complementary Multi-Encoder Vision-Language Learning From CLIP to DINO: Visual Encoders Shout in Multi-modal Large Language Models

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-13T20:33:17.149321Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-13T20:28:30.864143Z digest=sha256:0db4ba6b5b358a1686693b26b1e4cd8dc1a762e44f2bf03800268180df23fbc6

Observation 639975b0-9fe4-424d-9db0-ddbfa385ee24 · inbound

HAWK: Head Importance-Aware Visual Token Pruning in Multimodal Models cites this paper.

HAWK: Head Importance-Aware Visual Token Pruning in Multimodal Models From CLIP to DINO: Visual Encoders Shout in Multi-modal Large Language Models

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-11T07:01:05.567231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T17:20:19.653806Z digest=sha256:eb2f03ae80d4f4f4965ffe388604d2cde59d8a2c7c097900049638b1bea4cbcf

Observation a3803d08-b0de-4b1f-a851-f4046f935b27 · inbound

G-MIXER: Geodesic Mixup-based Implicit Semantic Expansion and Explicit Semantic Re-ranking for Zero-Shot Composed Image Retrieval cites this paper.

G-MIXER: Geodesic Mixup-based Implicit Semantic Expansion and Explicit Semantic Re-ranking for Zero-Shot Composed Image Retrieval From CLIP to DINO: Visual Encoders Shout in Multi-modal Large Language Models

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-10T11:50:19.964062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T11:50:00.850857Z digest=sha256:86075aa343f3751b252e424b484401cf257a18c094443f46cc6772b0e0bb151d

Observation c38471b3-d21d-4cb4-abf8-5df62635316d · inbound

Modeling Multi-Dimensional Cognitive States in Large Language Models under Cognitive Crowding cites this paper.

Modeling Multi-Dimensional Cognitive States in Large Language Models under Cognitive Crowding From CLIP to DINO: Visual Encoders Shout in Multi-modal Large Language Models

Reference 45

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T07:01:49.307858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-10T06:58:19.094492Z digest=sha256:1a36cc54ab1548c4b4a74ff52cd0e6974aa72a3230610f5634c178f261fbc6f8

Observation 3680a95e-b128-46bf-88c7-600e96d7739d · inbound

Are Natural-Domain Foundation Models Effective for Accelerated Cardiac MRI Reconstruction? cites this paper.

Are Natural-Domain Foundation Models Effective for Accelerated Cardiac MRI Reconstruction? From CLIP to DINO: Visual Encoders Shout in Multi-modal Large Language Models

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:26:11.344431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-08T09:08:28.326598Z digest=sha256:f612a52a9b03562c7670ac25e316b62b6607079bad58c355d9fc55aed5ab9acb

Observation 58db8de2-4657-4f0c-b486-b091a2229339 · inbound

MultiBreak: A Scalable and Diverse Multi-turn Jailbreak Benchmark for Evaluating LLM Safety cites this paper.

MultiBreak: A Scalable and Diverse Multi-turn Jailbreak Benchmark for Evaluating LLM Safety From CLIP to DINO: Visual Encoders Shout in Multi-modal Large Language Models

Reference 34

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T09:31:00.953629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-10T16:00:32.413225Z digest=sha256:08bb010cc2fc4d56154c6cd1f0b5852a7a667eb20246034945f621dfadf0c65d

Observation bcc9241a-18f3-4ff0-848f-ec5f2df21fe7 · inbound

PASA: A Principled Embedding-Space Watermarking Approach for LLM-Generated Text under Semantic-Invariant Attacks cites this paper.

PASA: A Principled Embedding-Space Watermarking Approach for LLM-Generated Text under Semantic-Invariant Attacks From CLIP to DINO: Visual Encoders Shout in Multi-modal Large Language Models

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:17:02.333089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-13T01:13:55.067785Z digest=sha256:0f9329fedb26ba2f02a5e8f7859cfd397a4d0a5060360f08520a889ca333b8e4

Observation 02cfb42a-25d1-4b9b-bb4f-5c3f5c1a48a5 · inbound

PASA: A Principled Embedding-Space Watermarking Approach for LLM-Generated Text under Semantic-Invariant Attacks cites this paper.

PASA: A Principled Embedding-Space Watermarking Approach for LLM-Generated Text under Semantic-Invariant Attacks From CLIP to DINO: Visual Encoders Shout in Multi-modal Large Language Models

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-06-30T23:45:08.405020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-06-30T23:37:21.982973Z digest=sha256:e7a44245217ab2ae0cc566f389c88b14a225ba7e1c33a86e194a2fe02a4d0b00

Observation 507e19a5-b352-4bdb-a469-82e92d246550 · inbound

New Wide-Net-Casting Jailbreak Attacks Risk Large Models cites this paper.

New Wide-Net-Casting Jailbreak Attacks Risk Large Models From CLIP to DINO: Visual Encoders Shout in Multi-modal Large Language Models

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-20T14:48:23.297324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T14:46:52.070071Z digest=sha256:44fdab0caa3eab81c6f123e15d1e26e3fd1bc20c10e22e6833ade75d71af17ee

Observation 3c2bf51d-b6ca-4524-8d43-f8951ff71cf6 · inbound

Mechanisms of Object Localization in Vision-Language Models cites this paper.

Mechanisms of Object Localization in Vision-Language Models From CLIP to DINO: Visual Encoders Shout in Multi-modal Large Language Models

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-20T06:48:05.803191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T06:46:20.141227Z digest=sha256:d7f78612889d5a453c735f916100ce6ffd30a902344d951b3af6f79565b383c2

Observation fdd383ec-bbde-4b1a-965e-ad04a6bb396d · inbound

Reducing Object Hallucination in LVLMs via Emphasizing Image-negative Tokens cites this paper.

Reducing Object Hallucination in LVLMs via Emphasizing Image-negative Tokens From CLIP to DINO: Visual Encoders Shout in Multi-modal Large Language Models

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-21T05:09:38.510067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-21T05:07:12.965100Z digest=sha256:11269524f04b76c94e34429bc7378ba8f0e324e69cbd635f47864d143193039c

Observation 817649ba-9fa7-452b-af17-eb8b87251fd9 · inbound

VaaWIT: Visual-Aware Adaptation of Large Language Models for Multilingual Web Image Translation cites this paper.

VaaWIT: Visual-Aware Adaptation of Large Language Models for Multilingual Web Image Translation From CLIP to DINO: Visual Encoders Shout in Multi-modal Large Language Models

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-06-30T13:34:40.271586Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-30T13:33:12.116833Z digest=sha256:cc358061423a40bfb24483ade8c13c801da2260827c8c594015839623220dd4f

Observation 55e98e71-8362-4592-b03c-286675d156b2 · inbound

SlotGCG: Exploiting the Positional Vulnerability in LLMs for Jailbreak Attacks cites this paper.

SlotGCG: Exploiting the Positional Vulnerability in LLMs for Jailbreak Attacks From CLIP to DINO: Visual Encoders Shout in Multi-modal Large Language Models

Reference 56

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T13:26:59.346516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-06-28T01:16:07.252429Z digest=sha256:30be899d7c814b299fcbbab25e58a18be36b6639f6d4f91512319508f2a6c2d7

Observation 2687d8ab-1295-47af-b1ca-0fb12f8f773d · inbound

Investigating The Security of Modern AI and Cloud Infrastructure cites this paper.

Investigating The Security of Modern AI and Cloud Infrastructure From CLIP to DINO: Visual Encoders Shout in Multi-modal Large Language Models

Reference 68

Resolution
verified exact
arxiv_id, observed 2026-07-04T08:29:42.327590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-26T11:31:39.910784Z digest=sha256:25bb3d4bf2d8c886aa6c431209b478390ea653018edf36c2295fb9498033172d

Observation 4ad0ef00-57c0-48b7-bbd6-c519d2da434e · inbound

EXPO-SQL: Execution-based Clause-level Policy Optimization for Text-to-SQL cites this paper.

EXPO-SQL: Execution-based Clause-level Policy Optimization for Text-to-SQL From CLIP to DINO: Visual Encoders Shout in Multi-modal Large Language Models

Reference 64

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T09:05:36.899599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-07-01T08:57:37.907354Z digest=sha256:84103a5c09fa772f455bd1f423764d00453975b7605c505c0f130670f2fbb18e

Observation a31b896d-ed4e-4428-8bb3-3714b9e5800b · inbound

CompressKV: Semantic-Retrieval-Guided KV-Cache Compression for Resource-Efficient Long-Context LLM Inference cites this paper.

CompressKV: Semantic-Retrieval-Guided KV-Cache Compression for Resource-Efficient Long-Context LLM Inference From CLIP to DINO: Visual Encoders Shout in Multi-modal Large Language Models

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-07-04T16:49:58.026479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-26T00:11:57.763089Z digest=sha256:9749a00e2752b9b2c340d7166eb3d27fa84b60bd9f80dcf34c01d57bf7b28c78

Observation 309ad413-30c8-43a7-b124-7c7ef35c92d5 · inbound

Mitigating Visual Degradation in MLLMs via Spatial-Spectral Visual Anchor Learning cites this paper.

Mitigating Visual Degradation in MLLMs via Spatial-Spectral Visual Anchor Learning From CLIP to DINO: Visual Encoders Shout in Multi-modal Large Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T15:10:13.086299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:10:13.086299Z digest=sha256:14dbb7cb5578780967284c356c3044a7bf9ac9f1c3954f86ddc12fd8164a7a99