Pith. sign in

Paper Citation Record · LEDGER

How Many Unicorns Are in This Image? A Safety Evaluation Benchmark for Vision LLMs

As of 11 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 18 inbound Pith citation observations for arXiv:2311.16101.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2311.16101 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 18 of 18 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 18 of 18 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-10T22:28:32.687691Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-17T22:30:23.249867Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 827bad60-ca14-4a9e-ad10-4911386eaf80 · inbound

Aligning Modalities in Vision Large Language Models via Preference Fine-tuning cites this paper.

Aligning Modalities in Vision Large Language Models via Preference Fine-tuning How Many Unicorns Are in This Image? A Safety Evaluation Benchmark for Vision LLMs

Reference 176

Resolution
verified exact
arxiv_id, observed 2026-05-17T10:58:53.471175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-17T10:58:53.215887Z digest=sha256:c996ab61da1cfcee9ccbd583686e9bacdb6cb61ed9847a19fb489e6edcabfb51

Observation 94f13af7-92c4-41cb-a373-51555c9858d8 · inbound

Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs cites this paper.

Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs How Many Unicorns Are in This Image? A Safety Evaluation Benchmark for Vision LLMs

Reference 130

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:05:03.754666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T00:05:03.547664Z digest=sha256:d22cb1dfff1a28468f708fe65d2251c7bda520f3530323dd79dcdd1a477a9a8e

Observation d1611879-a4cb-4d9c-8bf9-d70238c93b76 · inbound

MetaMorph: Multimodal Understanding and Generation via Instruction Tuning cites this paper.

MetaMorph: Multimodal Understanding and Generation via Instruction Tuning How Many Unicorns Are in This Image? A Safety Evaluation Benchmark for Vision LLMs

Reference 95

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T07:51:13.080598Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-17T07:51:12.953777Z digest=sha256:8ab51665e0e4e3bdf084d30e8cef01546d3a858585be899b9fef2796833d7b8f

Observation ca6e5ab2-adfa-469a-988e-4521c5096fdb · inbound

Spot Risks Before Speaking! Unraveling Safety Attention Heads in Large Vision-Language Models cites this paper.

Spot Risks Before Speaking! Unraveling Safety Attention Heads in Large Vision-Language Models How Many Unicorns Are in This Image? A Safety Evaluation Benchmark for Vision LLMs

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-10T22:28:32.687691Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:28:32.687691Z digest=sha256:87b1147587a19fee31dc258efb4cfc8cc172183d0ee5261f930e65cfa35f886b

Observation 40e8c8b3-92f1-475e-a9a8-ae3d45c4b630 · inbound

Are VLMs Ready for Autonomous Driving? An Empirical Study from the Reliability, Data, and Metric Perspectives cites this paper.

Are VLMs Ready for Autonomous Driving? An Empirical Study from the Reliability, Data, and Metric Perspectives How Many Unicorns Are in This Image? A Safety Evaluation Benchmark for Vision LLMs

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-10T21:44:55.952096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:44:55.952096Z digest=sha256:6ebebebf0970d8cbd2c4dd93b249f61cdee329912ff0b6a06e101ade32ce51c5

Observation 93b7c59c-a23f-4c39-8500-3867b5c5e602 · inbound

Eagle 2: Building Post-Training Data Strategies from Scratch for Frontier Vision-Language Models cites this paper.

Eagle 2: Building Post-Training Data Strategies from Scratch for Frontier Vision-Language Models How Many Unicorns Are in This Image? A Safety Evaluation Benchmark for Vision LLMs

Reference 122

Resolution
unresolved
no resolver link, observed 2026-08-10T18:04:34.480960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:04:34.480960Z digest=sha256:73eb1393273a08df1ea517913d13d151906b41eefc6146033933e9ee0a19ef79

Observation a058e217-c823-40c3-8f3d-92fbd41a2206 · inbound

A Survey of Safety on Large Vision-Language Models: Attacks, Defenses and Evaluations cites this paper.

A Survey of Safety on Large Vision-Language Models: Attacks, Defenses and Evaluations How Many Unicorns Are in This Image? A Safety Evaluation Benchmark for Vision LLMs

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T19:45:18.754300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T19:45:18.754300Z digest=sha256:5a5384533422f29d4f24ec73b481709dd7123a4eac98dd932887158bc2ab207c

Observation f3ba0300-05e8-43ea-a7e6-2f2c36a0b7ca · inbound

Retrieval Visual Contrastive Decoding to Mitigate Object Hallucinations in Large Vision-Language Models cites this paper.

Retrieval Visual Contrastive Decoding to Mitigate Object Hallucinations in Large Vision-Language Models How Many Unicorns Are in This Image? A Safety Evaluation Benchmark for Vision LLMs

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T13:57:07.419635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:57:07.419635Z digest=sha256:85fa1f1d32610f717700b6e1a01df106afa47c76f654de9fd100772e1f423234

Observation cf22a302-1e01-4e57-b838-dfdc96986a09 · inbound

Adversarial Attacks against Closed-Source MLLMs via Feature Optimal Alignment cites this paper.

Adversarial Attacks against Closed-Source MLLMs via Feature Optimal Alignment How Many Unicorns Are in This Image? A Safety Evaluation Benchmark for Vision LLMs

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T13:35:55.414755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:35:55.414755Z digest=sha256:3b49a2fb641efef0719c861583868f6f3dbe8b1d6f0cc3ad90cc32ff5b913889

Observation 1567cb64-eba7-4cfd-a546-41f62aa51bc2 · inbound

D-Fusion: Direct Preference Optimization for Aligning Diffusion Models with Visually Consistent Samples cites this paper.

D-Fusion: Direct Preference Optimization for Aligning Diffusion Models with Visually Consistent Samples How Many Unicorns Are in This Image? A Safety Evaluation Benchmark for Vision LLMs

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T13:23:11.230813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:23:11.230813Z digest=sha256:2e7d4e8f8d71d4ca2fddfd1ea676a801162700578e230fe6b45dcff040de0ed8

Observation 448ccc0f-a26a-4c91-ab4a-9b4750d221e9 · inbound

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments cites this paper.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments How Many Unicorns Are in This Image? A Safety Evaluation Benchmark for Vision LLMs

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:33.433505Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:33.433505Z digest=sha256:d3c9b533aa8b77828c956d19b4b1936dbe9eab2a3d6b2ad4656543211c885565

Observation 92032bf7-2422-4d67-8b64-828035486beb · inbound

The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs cites this paper.

The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs How Many Unicorns Are in This Image? A Safety Evaluation Benchmark for Vision LLMs

Reference 210

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:27.240375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:27.240375Z digest=sha256:8d0908f63d8718fb030c0f25524deedde65eea56c522b97186ffc515a40bcbf1

Observation ed55307c-ca35-43ce-8faf-54be1420f9b4 · inbound

One Object, Multiple Lies: A Benchmark for Cross-task Adversarial Attack on Unified Vision-Language Models cites this paper.

One Object, Multiple Lies: A Benchmark for Cross-task Adversarial Attack on Unified Vision-Language Models How Many Unicorns Are in This Image? A Safety Evaluation Benchmark for Vision LLMs

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T18:40:55.324196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:40:55.324196Z digest=sha256:4cf9d97d03d2a5eb521c51a8fe710f717506eda3d5424e05a629a7e8250fb9e6

Observation b049805a-607b-4e8f-b383-367b4963bad0 · inbound

OutSafe-Bench: A Benchmark for Multimodal Offensive Content Detection in Large Language Models cites this paper.

OutSafe-Bench: A Benchmark for Multimodal Offensive Content Detection in Large Language Models How Many Unicorns Are in This Image? A Safety Evaluation Benchmark for Vision LLMs

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-17T22:30:23.252260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T22:29:36.960961Z digest=sha256:1dba7494e0c6178b4f680ddef444ae8ad7d8bbfcf4fd6c610b0e66edec7eb1ac

Observation bdf9c0a4-a820-4468-9c70-e0f4b29403ba · inbound

Twins: Learn to Predict Unified Representations with Focal Loss cites this paper.

Twins: Learn to Predict Unified Representations with Focal Loss How Many Unicorns Are in This Image? A Safety Evaluation Benchmark for Vision LLMs

Reference 289

Resolution
unresolved
no resolver link, observed 2026-08-01T04:30:12.171419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T04:30:12.171419Z digest=sha256:fefe2af61226cdb34db3d5590f982dd64804548ec6fd45780e013bf1102cbe17

Observation af48c296-9a76-4654-be1d-b1b0e5752f51 · inbound

Two Sides of the Same Coin: Co-Evolving Search for Cross-Task Attacks on Vision-Language Models cites this paper.

Two Sides of the Same Coin: Co-Evolving Search for Cross-Task Attacks on Vision-Language Models How Many Unicorns Are in This Image? A Safety Evaluation Benchmark for Vision LLMs

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-04T13:51:18.171740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T13:51:18.171740Z digest=sha256:d32d752de87222bc1cbb36f2768e6f44c59d31d1cf80dfa7f8945bd7062f658f

Observation 2cc8404a-8f16-4ec7-8f39-8bc631564e22 · inbound

PRISM: Priority-aware Rubric Internalization via Structured Multimodal Data Synthesis cites this paper.

PRISM: Priority-aware Rubric Internalization via Structured Multimodal Data Synthesis How Many Unicorns Are in This Image? A Safety Evaluation Benchmark for Vision LLMs

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-08T17:13:25.940793Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:13:25.940793Z digest=sha256:07dafa23be486740257142410699cfd714db9fe3ab031e2632c83892de3539bf

Observation fd14ce9a-e180-4116-890f-78fce2594889 · inbound

PRISM: Priority-aware Rubric Internalization via Structured Multimodal Data Synthesis cites this paper.

PRISM: Priority-aware Rubric Internalization via Structured Multimodal Data Synthesis How Many Unicorns Are in This Image? A Safety Evaluation Benchmark for Vision LLMs

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T04:31:25.704790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:31:25.704790Z digest=sha256:40f21ecadbf12522b6c30832b04e27d013f3ac93ceba63605403fdb746ab4053