Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 18 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 43 inbound Pith citation observations for arXiv:2310.08825.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-15T15:10:13.086299Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T16:49:58.025162Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation c5cea414-c347-43ed-8d89-640ad01738a2 · inbound
Advancing Fine-Grained Visual Understanding with Multi-Scale Alignment in Multi-Modal Models From CLIP to DINO: Visual Encoders Shout in Multi-modal Large Language Models
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 87f820bf-13af-4c63-9f8a-c55ea640ef73 · inbound
Libra: Leveraging Temporal Images for Biomedical Radiology Analysis From CLIP to DINO: Visual Encoders Shout in Multi-modal Large Language Models
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d7a15a1a-c8b9-4686-adfe-973c94009a05 · inbound
VidHalluc: Evaluating Temporal Hallucinations in Multimodal Large Language Models for Video Understanding From CLIP to DINO: Visual Encoders Shout in Multi-modal Large Language Models
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b45d62d5-a49d-40b9-925b-a6e2f82afa10 · inbound
CPath-Omni: A Unified Multimodal Foundation Model for Patch and Whole Slide Image Analysis in Computational Pathology From CLIP to DINO: Visual Encoders Shout in Multi-modal Large Language Models
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f191937e-84f9-4da9-bf32-adca4d8e3361 · inbound
ComprehendEdit: A Comprehensive Dataset and Evaluation Framework for Multimodal Knowledge Editing From CLIP to DINO: Visual Encoders Shout in Multi-modal Large Language Models
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4bd58b9c-39b1-48e2-ad57-c9350f2b7e0d · inbound
Interpretable Face Anti-Spoofing: Enhancing Generalization with Multimodal Large Language Models From CLIP to DINO: Visual Encoders Shout in Multi-modal Large Language Models
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c7b5b60e-23e2-47cf-953a-30ef807b3ed0 · inbound
Diffusion Instruction Tuning From CLIP to DINO: Visual Encoders Shout in Multi-modal Large Language Models
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1b2c9241-8d5b-4c41-983b-dcd4728b5c64 · inbound
Toward Generalizable Forgery Detection and Reasoning From CLIP to DINO: Visual Encoders Shout in Multi-modal Large Language Models
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 4372cebe-1cb5-48a4-8acd-b28cb473a856 · inbound
RA-RRG: Multimodal Retrieval-Augmented Radiology Report Generation with Key Phrase Extraction From CLIP to DINO: Visual Encoders Shout in Multi-modal Large Language Models
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation dbba5313-2aad-4768-8d8a-e44aa9ffa657 · inbound
Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models From CLIP to DINO: Visual Encoders Shout in Multi-modal Large Language Models
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ea4cd2ef-2034-4a1b-8b7a-f9b908f0b8ea · inbound
Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor From CLIP to DINO: Visual Encoders Shout in Multi-modal Large Language Models
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 31d9f30d-3254-4f1c-8878-c08df639758a · inbound
METEOR: Multi-Encoder Collaborative Token Pruning for Efficient Vision Language Models From CLIP to DINO: Visual Encoders Shout in Multi-modal Large Language Models
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6d5705cf-9e95-4c13-baa3-b7bdb6ecc9bd · inbound
CompressKV: Semantic Retrieval Heads Know What Tokens are Not Important Before Generation From CLIP to DINO: Visual Encoders Shout in Multi-modal Large Language Models
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7de859b7-0815-4f8a-b1e1-53695a0ea7af · inbound
HEAL: A Hypothesis-Based Preference-Aware Analysis Framework From CLIP to DINO: Visual Encoders Shout in Multi-modal Large Language Models
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6c61d66c-4586-4d15-846f-be83e0498925 · inbound
Decoding Memories: An Efficient Pipeline for Self-Consistency Hallucination Detection From CLIP to DINO: Visual Encoders Shout in Multi-modal Large Language Models
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a92a6dea-f979-4319-b35a-d63074e2910a · inbound
TMUAD: Enhancing Logical Capabilities in Unified Anomaly Detection Models with a Text Memory Bank From CLIP to DINO: Visual Encoders Shout in Multi-modal Large Language Models
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a54931fd-ea48-473d-a5de-b2c9e84bb2d0 · inbound
Testing for LLM response differences: the case of a composite null consisting of semantically irrelevant query perturbations From CLIP to DINO: Visual Encoders Shout in Multi-modal Large Language Models
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 41b63a2f-a473-4cd1-a457-f6f7c807a93b · inbound
NP-LoRA: Null Space Projection for Subject-Style LoRA Fusion From CLIP to DINO: Visual Encoders Shout in Multi-modal Large Language Models
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation da0fb4f2-cd52-41d2-a5f9-f584920f3438 · inbound
VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs From CLIP to DINO: Visual Encoders Shout in Multi-modal Large Language Models
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c2906415-2ba2-4c97-b381-bcdb3c75e281 · inbound
Neuro-Symbolic Control with Large Language Models for Language-Guided Spatial Tasks From CLIP to DINO: Visual Encoders Shout in Multi-modal Large Language Models
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 008c6eab-522e-4181-8713-264b2bb71c45 · inbound
Who Endorsed It? Measuring Authority Bias Across Expertise Levels in Language Models From CLIP to DINO: Visual Encoders Shout in Multi-modal Large Language Models
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f10a2e3c-0f71-4694-8865-08eabaf7a67d · inbound
Harmonia: Algorithm-Hardware Co-Design for Memory- and Compute-Efficient BFP-based LLM Inference From CLIP to DINO: Visual Encoders Shout in Multi-modal Large Language Models
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7b8810df-b7bb-4a6a-85ec-1bd91ca6ba43 · inbound
Less is Enough: Synthesizing Diverse Data in LLM Feature Space with Sparse Autoencoders From CLIP to DINO: Visual Encoders Shout in Multi-modal Large Language Models
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bba09ca1-9f4f-42cc-8fde-16e58a55610e · inbound
InCoM: Intent-Driven Perception and Structured Coordination for Mobile Manipulation From CLIP to DINO: Visual Encoders Shout in Multi-modal Large Language Models
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 14df1184-58e6-4ca9-a7e3-8590e06f4864 · inbound
InCoM: Intent-Driven Perception and Structured Coordination for Mobile Manipulation From CLIP to DINO: Visual Encoders Shout in Multi-modal Large Language Models
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cfd77212-2fac-4242-bb44-7c4465f2cec7 · inbound
SpatialStack: Layered Geometry-Language Fusion for 3D VLM Spatial Reasoning From CLIP to DINO: Visual Encoders Shout in Multi-modal Large Language Models
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 6ab8a71c-97e8-40cf-87fd-de9bf6a49861 · inbound
CoME-VL: Scaling Complementary Multi-Encoder Vision-Language Learning From CLIP to DINO: Visual Encoders Shout in Multi-modal Large Language Models
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 639975b0-9fe4-424d-9db0-ddbfa385ee24 · inbound
HAWK: Head Importance-Aware Visual Token Pruning in Multimodal Models From CLIP to DINO: Visual Encoders Shout in Multi-modal Large Language Models
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation a3803d08-b0de-4b1f-a851-f4046f935b27 · inbound
G-MIXER: Geodesic Mixup-based Implicit Semantic Expansion and Explicit Semantic Re-ranking for Zero-Shot Composed Image Retrieval From CLIP to DINO: Visual Encoders Shout in Multi-modal Large Language Models
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation c38471b3-d21d-4cb4-abf8-5df62635316d · inbound
Modeling Multi-Dimensional Cognitive States in Large Language Models under Cognitive Crowding From CLIP to DINO: Visual Encoders Shout in Multi-modal Large Language Models
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 3680a95e-b128-46bf-88c7-600e96d7739d · inbound
Are Natural-Domain Foundation Models Effective for Accelerated Cardiac MRI Reconstruction? From CLIP to DINO: Visual Encoders Shout in Multi-modal Large Language Models
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 58db8de2-4657-4f0c-b486-b091a2229339 · inbound
MultiBreak: A Scalable and Diverse Multi-turn Jailbreak Benchmark for Evaluating LLM Safety From CLIP to DINO: Visual Encoders Shout in Multi-modal Large Language Models
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation bcc9241a-18f3-4ff0-848f-ec5f2df21fe7 · inbound
PASA: A Principled Embedding-Space Watermarking Approach for LLM-Generated Text under Semantic-Invariant Attacks From CLIP to DINO: Visual Encoders Shout in Multi-modal Large Language Models
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 02cfb42a-25d1-4b9b-bb4f-5c3f5c1a48a5 · inbound
PASA: A Principled Embedding-Space Watermarking Approach for LLM-Generated Text under Semantic-Invariant Attacks From CLIP to DINO: Visual Encoders Shout in Multi-modal Large Language Models
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 507e19a5-b352-4bdb-a469-82e92d246550 · inbound
New Wide-Net-Casting Jailbreak Attacks Risk Large Models From CLIP to DINO: Visual Encoders Shout in Multi-modal Large Language Models
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 3c2bf51d-b6ca-4524-8d43-f8951ff71cf6 · inbound
Mechanisms of Object Localization in Vision-Language Models From CLIP to DINO: Visual Encoders Shout in Multi-modal Large Language Models
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation fdd383ec-bbde-4b1a-965e-ad04a6bb396d · inbound
Reducing Object Hallucination in LVLMs via Emphasizing Image-negative Tokens From CLIP to DINO: Visual Encoders Shout in Multi-modal Large Language Models
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 817649ba-9fa7-452b-af17-eb8b87251fd9 · inbound
VaaWIT: Visual-Aware Adaptation of Large Language Models for Multilingual Web Image Translation From CLIP to DINO: Visual Encoders Shout in Multi-modal Large Language Models
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 55e98e71-8362-4592-b03c-286675d156b2 · inbound
SlotGCG: Exploiting the Positional Vulnerability in LLMs for Jailbreak Attacks From CLIP to DINO: Visual Encoders Shout in Multi-modal Large Language Models
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 2687d8ab-1295-47af-b1ca-0fb12f8f773d · inbound
Investigating The Security of Modern AI and Cloud Infrastructure From CLIP to DINO: Visual Encoders Shout in Multi-modal Large Language Models
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 4ad0ef00-57c0-48b7-bbd6-c519d2da434e · inbound
EXPO-SQL: Execution-based Clause-level Policy Optimization for Text-to-SQL From CLIP to DINO: Visual Encoders Shout in Multi-modal Large Language Models
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation a31b896d-ed4e-4428-8bb3-3714b9e5800b · inbound
CompressKV: Semantic-Retrieval-Guided KV-Cache Compression for Resource-Efficient Long-Context LLM Inference From CLIP to DINO: Visual Encoders Shout in Multi-modal Large Language Models
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 309ad413-30c8-43a7-b124-7c7ef35c92d5 · inbound
Mitigating Visual Degradation in MLLMs via Spatial-Spectral Visual Anchor Learning From CLIP to DINO: Visual Encoders Shout in Multi-modal Large Language Models
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.