Pith. sign in

Paper Citation Record · LEDGER

Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics

As of 18 August 2026, this Paper Citation Record lists 59 of 59 outbound references and 1 inbound Pith citation observation for arXiv:2412.10594.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.10594 v1

Coverage vector

measured 59 of 59 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T15:52:20.853170Z

measured 60 of 60 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-12T00:53:35.188721Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-12T08:36:26.482689Z

Reference resolution

59 of 59 outbound references displayed

  • verified exact0
  • verified fuzzy30
  • unresolved29
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 17c76d01-0a43-4413-9ca1-1c11c3c02838 · outbound

This paper cites GPT-4 Technical Report.

Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T15:52:20.513627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:52:20.513627Z digest=sha256:9bda16e899adedc1211537ec66a291985d063c0922e8130644a2c099b3ec8f85

Observation 8a9ada32-5c20-4329-ae7a-bb6c736e9d9f · outbound

This paper cites Getting vit in shape: Scaling laws for compute-optimal model design.

Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics Getting vit in shape: Scaling laws for compute-optimal model design

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:52:22.028682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T15:52:20.519143Z digest=sha256:13b6a2b302f05daea293c529ecc0307c53df5c5692dda11dc7f021162600bc76

Observation fcbf14c3-cec0-4f98-9c87-b6bc37f76105 · outbound

This paper cites Improving image captioning descriptive- ness by ranking and llm-based fusion.

Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics Improving image captioning descriptive- ness by ranking and llm-based fusion

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T15:52:20.524263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:52:20.524263Z digest=sha256:375dcdb103e1da7f5487a9e515cdd548f5d9bfd3c7231704929a07878195656c

Observation c51dfb68-7948-473f-9ac6-d97d834435e5 · outbound

This paper cites Learning a deep single image contrast enhancer from multi-exposure images.

Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics Learning a deep single image contrast enhancer from multi-exposure images

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T15:52:20.528675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:52:20.528675Z digest=sha256:9ea8b63b82a1b0d2eb7d8521799dd81e3059d97d6cbec239ccdfff3b594fd341

Observation 77e17d57-450e-4094-9598-a01932b70e4d · outbound

This paper cites Emerg- ing properties in self-supervised vision transformers.

Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics Emerg- ing properties in self-supervised vision transformers

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:52:21.991793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T15:52:20.532907Z digest=sha256:e3cb5f6d451b101edf1ceaeb5d6f20d761422f961e94795433c1a512509abe25

Observation 46faf46c-e022-4296-a3ee-a7738fa3cfaf · outbound

This paper cites Reproducible scal- ing laws for contrastive language-image learning.

Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics Reproducible scal- ing laws for contrastive language-image learning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T15:52:20.537679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:52:20.537679Z digest=sha256:c6aebba24f81277e256930457a6a2df3f82877ec50b43cce702159ca3b3afa0e

Observation 5301a0e6-3010-4e86-bcfd-0c4e4172cdf5 · outbound

This paper cites Adversarially robust clip mod- els induce better (robust) perceptual metrics.

Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics Adversarially robust clip mod- els induce better (robust) perceptual metrics

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:52:21.957738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T15:52:20.544410Z digest=sha256:44c6372cfa611058082788d8b981054b30c2b45ab778c9be3d964bdcaec163e8

Observation 5f256352-438c-4a75-b6b7-5d347f933941 · outbound

This paper cites Dream- sim: Learning new dimensions of human visual similarity using synthetic data.

Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics Dream- sim: Learning new dimensions of human visual similarity using synthetic data

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:52:21.936354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T15:52:20.551039Z digest=sha256:04bfb6ac69b1517cffbc5ce048c2651ef9f2aa12efb110ca8bac7510915e4e8a

Observation 248a84da-d421-42b5-b827-57a691d103b1 · outbound

This paper cites BLINK: Multimodal Large Language Models Can See but Not Perceive.

Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T15:52:20.559022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:52:20.559022Z digest=sha256:c08dab203131c30ece642883cbcf2882bb5e8be34dfc482d61bb4d95ca342180

Observation 78ca037e-9fe7-4594-8b72-ead91024827c · outbound

This paper cites R-LPIPS: An adversarially robust perceptual similarity metric.

Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics R-LPIPS: An adversarially robust perceptual similarity metric

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:52:21.921279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T15:52:20.564908Z digest=sha256:c443ca12b37c8d65ccc4bcd328704820e77cfd11a3957956883290651135d8b7

Observation 7553f00d-4b45-46b2-a213-2e46d7755918 · outbound

This paper cites EMMA: Efficient Visual Alignment in Multi-Modal LLMs.

Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics EMMA: Efficient Visual Alignment in Multi-Modal LLMs

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T15:52:20.570468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:52:20.570468Z digest=sha256:5cc781146ec6deb81b51e2727c96b378226dc1759725e6eb45e62e87822e8f9e

Observation 2f00ae9a-dee6-4749-a102-8b2fc039c5fc · outbound

This paper cites Lipsim: A provably robust perceptual similarity metric.

Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics Lipsim: A provably robust perceptual similarity metric

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:52:21.904307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T15:52:20.575211Z digest=sha256:5cc86ba87d55476da97c0b080bf3900fd8111ec864a5970703953d4446ca7335

Observation 055ed614-487c-4d54-a188-f7d8395924af · outbound

This paper cites Generative adversarial networks.

Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics Generative adversarial networks

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:52:21.888020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T15:52:20.579289Z digest=sha256:6f58bc0ac911fb949df0e283a2998dbbe3c562bc5504d724f45134538e64707e

Observation f0b984e3-ae3b-41a1-a946-365a82742d77 · outbound

This paper cites Masked autoencoders are scalable vision learners.

Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics Masked autoencoders are scalable vision learners

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:52:21.874994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T15:52:20.583662Z digest=sha256:002a87b4f64099652dbebbec609d2185b6006f19e5ce39d7c0803aa6a4c7c8af

Observation 768a2a18-9037-4942-9b8b-e99759ceb529 · outbound

This paper cites Clipscore: A reference-free evaluation met- ric for image captioning.

Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics Clipscore: A reference-free evaluation met- ric for image captioning

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:52:21.861345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T15:52:20.588191Z digest=sha256:68ddd684919742d0f4c16e2b5acf4876329b14b29c0fdff24e21f95602d4fce2

Observation 7d9e3fd5-4727-4813-983a-4c43c38a8296 · outbound

This paper cites Denoising dif- fusion probabilistic models.

Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics Denoising dif- fusion probabilistic models

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:52:21.845892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T15:52:20.593566Z digest=sha256:f23675826abb6014ac732057d6f6b15d2a17c34033b7c2578cfc3727a22b6871

Observation 4bca766e-af38-420e-b26e-12dec1e77cc4 · outbound

This paper cites Koniq-10k: An ecologically valid database for deep learning of blind image quality assessment.

Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics Koniq-10k: An ecologically valid database for deep learning of blind image quality assessment

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:52:21.829191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T15:52:20.598218Z digest=sha256:ad98f5641fd2a2a3fa87d7756448ac8ca2d3cf9584ff450e3031d7fc30c02d5d

Observation 7f102026-893e-49de-817b-2145504c12bb · outbound

This paper cites LoRA: Low-rank adaptation of large language models.

Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics LoRA: Low-rank adaptation of large language models

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:52:21.812630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T15:52:20.602776Z digest=sha256:cfcbe86c6224d3043b4ed460dfc12387ad8a01c0c1c34da28f54f7c642454480

Observation 0706775c-7b58-4056-854e-575a58f0be4a · outbound

This paper cites Aesexpert: Towards multi-modality foun- dation model for image aesthetics perception.

Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics Aesexpert: Towards multi-modality foun- dation model for image aesthetics perception

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:52:21.797672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T15:52:20.607404Z digest=sha256:0d43a9d27c58c8cbff378fa8defa3f5b72c44e89c6ffca884eae36d3234e1f4f

Observation ed87fb23-beca-42d8-b7e7-1e1386342417 · outbound

This paper cites HQ-Edit: A High-Quality Dataset for Instruction-based Image Editing.

Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics HQ-Edit: A High-Quality Dataset for Instruction-based Image Editing

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T15:52:20.612245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:52:20.612245Z digest=sha256:d64918af1ee548822b27c017c5cae5c3daf6b540286834582c0e193a78ecf663

Observation 2d7d4595-41c1-44fa-9c73-35d5c952ef8c · outbound

This paper cites MANTIS: Interleaved Multi-Image Instruction Tuning.

Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics MANTIS: Interleaved Multi-Image Instruction Tuning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T15:52:20.618332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:52:20.618332Z digest=sha256:b3d9f88adda371dabec44b97fdda05985d05c64de5c0bc41cac3e832c7a99d3a

Observation 87c04987-a20c-4660-bb21-60d9fa57363f · outbound

This paper cites Pipal: a large-scale image quality assessment dataset for perceptual image restoration.

Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics Pipal: a large-scale image quality assessment dataset for perceptual image restoration

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:52:21.782689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T15:52:20.623290Z digest=sha256:e7a59783afadee7614b255d044391f7325256deadb1c5caaed91de94135d0bce

Observation 6adcabd3-294e-4608-9c12-f8dbd0ee536a · outbound

This paper cites ImagenHub: Standardizing the evaluation of conditional image generation models.

Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics ImagenHub: Standardizing the evaluation of conditional image generation models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T15:52:20.628436Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:52:20.628436Z digest=sha256:94ceb10ca09ca4cfef054b50a18996e6017f94bbfe93fb578aa5e2f4594ecf82

Observation 87f42d74-cee5-4d53-8ad1-b622642bf1c1 · outbound

This paper cites UMIC: An Unreferenced Metric for Image Captioning via Contrastive Learning.

Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics UMIC: An Unreferenced Metric for Image Captioning via Contrastive Learning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T15:52:20.635404Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:52:20.635404Z digest=sha256:460dbfc146aa5c049b80739c6d19e55b96e912a381ff82f25f8c6a13a2b00857

Observation 0ebd9825-6b74-4da0-b548-49e7270fa039 · outbound

This paper cites Agiqa-3k: An open database for ai-generated image quality assessment.

Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics Agiqa-3k: An open database for ai-generated image quality assessment

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:52:21.766045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T15:52:20.640156Z digest=sha256:c1d236a47f875766fd35cbacb561406cdd5e51f8f88136f9635465ec04cbacda

Observation 38f8cb07-b66c-48cc-b277-01fadedc1078 · outbound

This paper cites AIGIQA-20K: A Large Database for AI-Generated Image Quality Assessment.

Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics AIGIQA-20K: A Large Database for AI-Generated Image Quality Assessment

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T15:52:20.645204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:52:20.645204Z digest=sha256:c827ee3aef52178c2c534385cd6ece42a7f9553a93d4f98ee0be6dbecc5ad1a0

Observation 15a2c559-3316-493e-be5f-3d9e2562158b · outbound

This paper cites LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models.

Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T15:52:20.651720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:52:20.651720Z digest=sha256:e5fce041cba836d1a47b4586a87275d24d828123221348b9e3d527fee6fa2a6e

Observation 75bf0df2-9641-4f84-b2b8-f4038de5a9b9 · outbound

This paper cites BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models.

Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T15:52:20.657956Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:52:20.657956Z digest=sha256:165f5598f521e1062ccdd7bae210fd1de0eeea4274187fdb8d00c899f3d0e5c4

Observation 8b6c9e92-3fe4-4e4f-9893-65b40f1222a9 · outbound

This paper cites Kadid-10k: A large-scale artificially distorted iqa database.

Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics Kadid-10k: A large-scale artificially distorted iqa database

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:52:21.747189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T15:52:20.663794Z digest=sha256:281dfd8edce9bc7108dbd94abfb028d684bc563e23eb236c506ddad8dbdbe026

Observation 2856e819-fd2f-4b73-b33a-94b9572dd042 · outbound

This paper cites Microsoft coco: Common objects in context.

Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics Microsoft coco: Common objects in context

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:52:21.731632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T15:52:20.668842Z digest=sha256:799a9b8fd1337c32069983682ec4912dcd55af1bbf449a66d26e0ed7744002ae

Observation 06dafb6f-09b3-47b4-ac97-9522d6245161 · outbound

This paper cites an unresolved cited work.

Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:52:21.713584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T15:52:20.674338Z digest=sha256:b1f732fda82e6a88d81d269e26e5bcf4364ed4ca07e0d670b78b510d7be66a8d

Observation f6262e8d-72fc-4613-89b3-5594fd668549 · outbound

This paper cites Improved baselines with visual instruction tuning.

Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics Improved baselines with visual instruction tuning

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:52:21.699013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T15:52:20.679260Z digest=sha256:ae97cc465dd3bfa73abf7c2536c5134773f99b058d0782af24fdfa23a349d740

Observation 290c40b1-d467-494b-8394-25b0777f7ba9 · outbound

This paper cites FuseDream: Training-Free Text-to-Image Generation with Improved CLIP+GAN Space Optimization.

Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics FuseDream: Training-Free Text-to-Image Generation with Improved CLIP+GAN Space Optimization

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T15:52:20.683958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:52:20.683958Z digest=sha256:6ff0a9bbd965d555721866431d7518d47498111fe0203529cf20a01863377fcd

Observation 6d45ef75-527d-43ad-a1f5-789ebdbca73f · outbound

This paper cites RoBERTa: A Robustly Optimized BERT Pretraining Approach.

Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T15:52:20.689029Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:52:20.689029Z digest=sha256:31090faa8dc95f017cf3cb62d5adc1a12c6fa69bb56846d8fc44477403cba876

Observation 1e678f20-634d-452a-97a4-631ed9541289 · outbound

This paper cites Vandermeulen, and Simon Kornblith.

Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics Vandermeulen, and Simon Kornblith

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:52:21.680734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T15:52:20.694008Z digest=sha256:a15eac6d9a0ceb7404c66cafa46f70d228068c60e4d97a3a7668b9a56abd4939

Observation 0bd7fc4f-a57b-4108-bc20-f44c8c52d153 · outbound

This paper cites Lost in quantization: Improving particu- lar object retrieval in large scale image databases.

Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics Lost in quantization: Improving particu- lar object retrieval in large scale image databases

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:52:21.665683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T15:52:20.699753Z digest=sha256:10b1e750576b8a531049c1aa93f27c4af89ccd041d7756ad31c82dd52dd2fab5

Observation 6dd01535-2934-4613-a92b-50ebe1af757c · outbound

This paper cites SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis.

Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T15:52:20.706284Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:52:20.706284Z digest=sha256:62c50f70ebc3d1d91f1eb368d2399315fbd5c6631afe3ae2646fa7ceb9520726

Observation 14160a5e-c550-463e-a728-50a63b09ce36 · outbound

This paper cites Pieapp: Perceptual image-error assessment through pairwise preference.

Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics Pieapp: Perceptual image-error assessment through pairwise preference

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:52:21.649485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T15:52:20.714728Z digest=sha256:cff6ebc7637bccf38eab38a880adbe1006a3ed52840f991e3bbde9b4eda0edae

Observation 7dc182ff-f0aa-4b51-b3f7-dcbc099e7735 · outbound

This paper cites Learning transferable visual models from natural language supervision.

Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics Learning transferable visual models from natural language supervision

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:52:21.634785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T15:52:20.719641Z digest=sha256:03401df9eaa040d29709ddd3ff3cd6697c99d0d990a813b70c2c5507caa9a035

Observation 6d3e3eb5-774f-4ef3-92a7-aa683fd2a1eb · outbound

This paper cites Positive-augmented contrastive learning for image and video captioning evaluation.

Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics Positive-augmented contrastive learning for image and video captioning evaluation

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:52:21.619423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T15:52:20.725645Z digest=sha256:54e8d7d34d3969b2ddefaf1aeb6ac2f5e9e94b65502f53b2978f3a4fa36032fc

Observation 986fc960-2fd1-42ca-acb7-081ec7faf1c3 · outbound

This paper cites When Does Perceptual Alignment Benefit Vision Representations?.

Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics When Does Perceptual Alignment Benefit Vision Representations?

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T15:52:20.730292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:52:20.730292Z digest=sha256:9c2f228f57ba3a2ec03311baf52fe000cf8634593b1a717d1da06eae0d3b444e

Observation 88bae7ca-aa88-4e61-96b3-2d5eb9dea848 · outbound

This paper cites Chameleon: Mixed-Modal Early-Fusion Foundation Models.

Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics Chameleon: Mixed-Modal Early-Fusion Foundation Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-11T15:52:20.735706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:52:20.735706Z digest=sha256:cab41fefecfb8c3d92de673241aab71afc7d34d6d8ee64ae5f5911129094fb17

Observation 9ae66f87-1dc5-4b7f-8485-d15526238d19 · outbound

This paper cites Polos: Multimodal metric learning from human feed- back for image captioning.

Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics Polos: Multimodal metric learning from human feed- back for image captioning

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:52:21.604954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T15:52:20.741709Z digest=sha256:af3cc675845238ea9b351f6756d14f7134160be81097dcedb5139b2a5026f66c

Observation ab87bdbb-d304-417f-b270-810d1ab4a970 · outbound

This paper cites MuirBench: A Comprehensive Benchmark for Robust Multi-image Understanding.

Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics MuirBench: A Comprehensive Benchmark for Robust Multi-image Understanding

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-11T15:52:20.746009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:52:20.746009Z digest=sha256:c4eae5795f4000e00896df8d76b12358fd1965b4a6f3fe88263b183137ae8cd8

Observation 240ca792-5f11-4660-926e-af9f714084f4 · outbound

This paper cites Ex- ploring clip for assessing the look and feel of images.

Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics Ex- ploring clip for assessing the look and feel of images

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:52:21.588807Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T15:52:20.751054Z digest=sha256:fea225fff0c752bafc261d7897b94ed7d746b156222bda7c7381a8d10d5d4f3d

Observation 6baf2c07-0a34-4d85-b17c-83b77d28dfcb · outbound

This paper cites Q-Bench: A Benchmark for General-Purpose Foundation Models on Low-level Vision.

Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics Q-Bench: A Benchmark for General-Purpose Foundation Models on Low-level Vision

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-11T15:52:20.757343Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:52:20.757343Z digest=sha256:188c03663f25f651880374a4bca19873610091f2b53963f7b557af3609e57594

Observation 53337deb-fa14-454c-bb61-b742d6f0d5e0 · outbound

This paper cites Towards Open-ended Visual Quality Comparison.

Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics Towards Open-ended Visual Quality Comparison

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-11T15:52:20.764418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:52:20.764418Z digest=sha256:3f3d924aaf43659cc956960d739653929a78e3edc9f4aa474b5aee36674aae38

Observation 83722df9-9853-4a9c-8b18-115afc89969f · outbound

This paper cites Human Preference Score v2: A Solid Benchmark for Evaluating Human Preferences of Text-to-Image Synthesis.

Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics Human Preference Score v2: A Solid Benchmark for Evaluating Human Preferences of Text-to-Image Synthesis

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-11T15:52:20.771725Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:52:20.771725Z digest=sha256:2acf23c9baaf2df0edf6618bbe23c80f3a20160809b83b02daec394ae77e1863

Observation 75e7f3ee-210f-47cf-a208-2aa100b3880f · outbound

This paper cites Imagere- ward: Learning and evaluating human preferences for text- to-image generation.

Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics Imagere- ward: Learning and evaluating human preferences for text- to-image generation

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:52:21.546559Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T15:52:20.783050Z digest=sha256:23ccde89840104277c002aed195d226c24373777341841d6848cd90192122f1c

Observation 16d4f2eb-1bb6-4fff-afe1-99534ca33c07 · outbound

This paper cites mplug- owl2: Revolutionizing multi-modal large language model with modality collaboration.

Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics mplug- owl2: Revolutionizing multi-modal large language model with modality collaboration

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:52:21.530751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T15:52:20.790872Z digest=sha256:3e213c923fd9b1c5410a60c0595d71d5e7b009dbe4fa8491ae246658b7c126e3

Observation 021d40ec-759e-4189-9059-8c0083d87534 · outbound

This paper cites Sigmoid Loss for Language Image Pre-Training.

Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics Sigmoid Loss for Language Image Pre-Training

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-11T15:52:20.798079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:52:20.798079Z digest=sha256:601fab41b50f93147ea595e8f304dec787d30d86270fde512306c67bbb010a4b

Observation 0109bb72-16dd-4d30-afd5-29dd22b2c998 · outbound

This paper cites Text-to-image Diffusion Models in Generative AI: A Survey.

Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics Text-to-image Diffusion Models in Generative AI: A Survey

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-11T15:52:20.803758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:52:20.803758Z digest=sha256:1d07a3e042919ce4611b5aee97342100a0bdde9f3be3d4221b2530a51887ee84

Observation 63550c03-bfd8-424b-a306-0f8b6a1c56ca · outbound

This paper cites Magicbrush: A manually annotated dataset for instruction- guided image editing.

Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics Magicbrush: A manually annotated dataset for instruction- guided image editing

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:52:21.513857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T15:52:20.812589Z digest=sha256:132dcdd76e25c31625f4b44e88184e9f6409c4c3fe90170bebe275386a0c9ff9

Observation fd1a6379-26b0-4115-86de-fda62eba9f85 · outbound

This paper cites Efros, Eli Shecht- man, and Oliver Wang.

Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics Efros, Eli Shecht- man, and Oliver Wang

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:52:21.496296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T15:52:20.823969Z digest=sha256:9418364c8113ce4e8492465b2238529a30883350503c14b0ff485df966628e73

Observation c7625820-c4d3-4a4b-808b-23313a459cfb · outbound

This paper cites Blind image quality assessment via vision- language correspondence: A multitask learning perspective.

Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics Blind image quality assessment via vision- language correspondence: A multitask learning perspective

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:52:21.480909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T15:52:20.831283Z digest=sha256:aafdccb7fab6060aeda9d688d933932d034a7f99f82b9856f54729d82fb577b7

Observation afb4e741-d43f-4779-8027-ba6b89d1f49b · outbound

This paper cites A-Bench: Are LMMs Masters at Evaluating AI-generated Images?.

Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics A-Bench: Are LMMs Masters at Evaluating AI-generated Images?

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-11T15:52:20.839190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:52:20.839190Z digest=sha256:5f9973ae87d22a99d73a1ee3aacebc65ef4a76ed4d8c9ff701d80f16ba4c02ab

Observation 6155f7f7-7f0b-45a1-b3da-ddc8328ad328 · outbound

This paper cites 2AFC Prompting of Large Multimodal Models for Image Quality Assessment.

Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics 2AFC Prompting of Large Multimodal Models for Image Quality Assessment

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-11T15:52:20.845040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:52:20.845040Z digest=sha256:db809394a5e7649bc6ebbff3e646467f93c84c800b0ef680722ccab45de5ae1e

Observation 8d03083d-7bac-494f-bcf0-dd6c7656c0fa · outbound

This paper cites Adaptive Image Quality Assessment via Teaching Large Multimodal Model to Compare.

Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics Adaptive Image Quality Assessment via Teaching Large Multimodal Model to Compare

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-11T15:52:20.853170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:52:20.853170Z digest=sha256:dcdc5d4312caf958bca85e3ed2160715a816f7d2622c8730698a6d9ba988c05c

Observation d86a91e8-aa96-4272-9426-cdbabaa8de2c · outbound

This paper cites an unresolved cited work.

Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics Unresolved cited work

Reference 2023

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:52:21.563678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T15:52:20.777233Z digest=sha256:a658448916de53f08a900a2dee0e6e955f162786bc081044fa33c9e25a10f691

Pith citing papers

Observation 199e40c4-055f-4d48-bd48-e116c72c03bf · inbound

SYNCR: A Cross-Video Reasoning Benchmark with Synthetic Grounding cites this paper.

SYNCR: A Cross-Video Reasoning Benchmark with Synthetic Grounding Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:36:26.485902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-12T00:53:35.188721Z digest=sha256:0f87be75b3b4654034bb09237150a35a46e32005c1730825faf27df418641a4c