Pith. sign in

Paper Citation Record · LEDGER

AVA-Bench: Atomic Visual Ability Benchmark for Vision Foundation Models

As of 6 August 2026, this Paper Citation Record lists 100 of 158 outbound references and 1 inbound Pith citation observation for arXiv:2506.09082.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.09082 v5

Coverage vector

measured 100 of 158 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-19T11:12:41.130806Z

measured 101 of 101 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-14T22:19:38.973091Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

100 of 158 outbound references displayed

  • verified exact33
  • verified fuzzy67
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 437a8071-a879-40b2-94a4-b357927457f6 · outbound

This paper cites GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints.

AVA-Bench: Atomic Visual Ability Benchmark for Vision Foundation Models GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-19T11:13:02.837360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T11:12:41.130806Z digest=sha256:68db6305a26b2adcd46fbacdaf51317e002d790eff9f87783f2e77185e1f45db

Observation b7803521-163e-40d4-9ffa-3c0623ce07ba · outbound

This paper cites Foundation models defining a new era in vision: a survey and outlook.

AVA-Bench: Atomic Visual Ability Benchmark for Vision Foundation Models Foundation models defining a new era in vision: a survey and outlook

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T11:13:03.052914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T11:12:41.130806Z digest=sha256:d533c0cf3b099cbc4facdd40c34f48c1ba8ae9ced97975e6dd3d0e98daeb9947

Observation 8286d9f7-720f-4b99-9349-4e1d4a0063a6 · outbound

This paper cites Eureka: Evaluating and Understanding Large Foundation Models.

AVA-Bench: Atomic Visual Ability Benchmark for Vision Foundation Models Eureka: Evaluating and Understanding Large Foundation Models

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-19T11:13:02.874659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T11:12:41.130806Z digest=sha256:8fd89cc98860b9a06d4eb30b5fda80cdabe04a784257b67747bae729f38583dd

Observation 08a5b9b2-fb20-4a66-9bb4-73927304f280 · outbound

This paper cites Scene text visual question answering.

AVA-Bench: Atomic Visual Ability Benchmark for Vision Foundation Models Scene text visual question answering

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T11:13:03.033241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T11:12:41.130806Z digest=sha256:df3f9c5d0e9e72b7710ffcc356e02908ed95c641133de7733c4662e9c524c3b2

Observation 68c006e2-0e69-42f3-909e-e0f73ecc247d · outbound

This paper cites On the Opportunities and Risks of Foundation Models.

AVA-Bench: Atomic Visual Ability Benchmark for Vision Foundation Models On the Opportunities and Risks of Foundation Models

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-19T11:13:02.890234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T11:12:41.130806Z digest=sha256:de4ca5c945a869145f5619cd0d2aad08c8475ae13f6b47cf195e10676b0f50f2

Observation 6ce07b22-b15f-4c6d-85eb-bba5f39b38c4 · outbound

This paper cites Emerging properties in self-supervised vision transformers.

AVA-Bench: Atomic Visual Ability Benchmark for Vision Foundation Models Emerging properties in self-supervised vision transformers

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T11:13:03.040320Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T11:12:41.130806Z digest=sha256:a4533ad5ffb463d56c803fee5878c36bc899cd0ab47d8a686dfa68c550d5bde8

Observation 3999625d-bc2c-457b-adac-093c94514c8e · outbound

This paper cites Decomposing complex visual comprehension into atomic visual skills for vision language models.

AVA-Bench: Atomic Visual Ability Benchmark for Vision Foundation Models Decomposing complex visual comprehension into atomic visual skills for vision language models

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T11:13:03.017518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T11:12:41.130806Z digest=sha256:ecf90383bd2140f6321becb8d63797e649d29a7241220cd2f88ccd917dd9a108

Observation 92cc5132-7175-4c94-a15a-9ce0adad6381 · outbound

This paper cites Sharegpt4v: Improving large multi-modal models with better captions.

AVA-Bench: Atomic Visual Ability Benchmark for Vision Foundation Models Sharegpt4v: Improving large multi-modal models with better captions

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T11:13:03.068609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T11:12:41.130806Z digest=sha256:1d876c28802e85b8d892ee1d789a9bd7a65602e6063b0403bfb98fd6306177e7

Observation 279d5f37-ea74-49f7-8507-b45140b68e69 · outbound

This paper cites Segment Anything Model (SAM) Enhanced Pseudo Labels for Weakly Supervised Semantic Segmentation.

AVA-Bench: Atomic Visual Ability Benchmark for Vision Foundation Models Segment Anything Model (SAM) Enhanced Pseudo Labels for Weakly Supervised Semantic Segmentation

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-19T11:13:02.896114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T11:12:41.130806Z digest=sha256:7f3625de4ffc3dae854347eef55534494852dcc51cb50144fba31360d0d1159c

Observation 720ed980-043f-4937-95ef-8c99458adeaa · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

AVA-Bench: Atomic Visual Ability Benchmark for Vision Foundation Models Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-19T11:13:02.812538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T11:12:41.130806Z digest=sha256:56d54672fc7f5606d511017912b01b9576d8b52fe8985d2e3a15d66eb459ebe6

Observation 6fd232ec-337c-4e50-9796-f460cee33860 · outbound

This paper cites On Domain-Adaptive Post-Training for Multimodal Large Language Models.

AVA-Bench: Atomic Visual Ability Benchmark for Vision Foundation Models On Domain-Adaptive Post-Training for Multimodal Large Language Models

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-19T11:13:02.828424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T11:12:41.130806Z digest=sha256:1ad6420cb7d12d01863930239826ee7d3ec449d2621e7d88047e1263f1b05c7c

Observation dd796da5-a3e7-4af0-afcd-154ddd1da6c1 · outbound

This paper cites MegaCOIN: Enhancing Medium-Grained Color Perception for Vision-Language Models.

AVA-Bench: Atomic Visual Ability Benchmark for Vision Foundation Models MegaCOIN: Enhancing Medium-Grained Color Perception for Vision-Language Models

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-19T11:13:02.869014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T11:12:41.130806Z digest=sha256:4633932a9a77fe5d8ec2f82f12e2b7edf158b69b97480c3d366fb0ffd3ab834b

Observation 348c2c32-8dab-46db-9e03-282ad73bc4f9 · outbound

This paper cites Palm: Scaling language modeling with pathways.

AVA-Bench: Atomic Visual Ability Benchmark for Vision Foundation Models Palm: Scaling language modeling with pathways

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T11:13:03.021070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T11:12:41.130806Z digest=sha256:b69ea1166b6f933c92ff629513e46909b52ee2debd269aed22743d9b201a455b

Observation e04d3eea-4b88-4095-b940-aa224aff8e83 · outbound

This paper cites Prompt-CAM: Making Vision Transformers Interpretable for Fine-Grained Analysis.

AVA-Bench: Atomic Visual Ability Benchmark for Vision Foundation Models Prompt-CAM: Making Vision Transformers Interpretable for Fine-Grained Analysis

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-19T11:13:02.882874Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T11:12:41.130806Z digest=sha256:1980c86fb811ad245559e8585bdc21ac2c28284b42728f99453bda8b44ab224a

Observation 4c968fa0-f649-4219-90fb-859a301ebe34 · outbound

This paper cites Cimpoi, S.

AVA-Bench: Atomic Visual Ability Benchmark for Vision Foundation Models Cimpoi, S

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T11:13:03.120546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T11:12:41.130806Z digest=sha256:6ffddc9a3a6c0036e69b05d3512bc52e4a7275798516c2fdb0204edf5cf21b72

Observation 114f337c-b6ee-4faf-9bbd-9834a1597742 · outbound

This paper cites Humanvlm: Foundation for human- scene vision-language model.

AVA-Bench: Atomic Visual Ability Benchmark for Vision Foundation Models Humanvlm: Foundation for human- scene vision-language model

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T11:13:03.115722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T11:12:41.130806Z digest=sha256:d81107fd9470710d185e580b784cc9f0869d840319b73a7a1223ee785dba7a77

Observation 0f0ea91a-202e-4131-8681-5cf6fd4cb4dd · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

AVA-Bench: Atomic Visual Ability Benchmark for Vision Foundation Models An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-05-19T11:13:02.885253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T11:12:41.130806Z digest=sha256:78e36ef322b60a40e2c9060725f1e4f660317e1c169c343a34c48aee8ae23a40

Observation 28701936-19fc-490e-a422-603c0e74bd02 · outbound

This paper cites Shape and texture recognition in large vision-language models.

AVA-Bench: Atomic Visual Ability Benchmark for Vision Foundation Models Shape and texture recognition in large vision-language models

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-19T11:13:02.857949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T11:12:41.130806Z digest=sha256:7fa1c273981f5e13fb03207ec417ab6856ff60a94c4afd8c78b8c654c2ae9156

Observation 4edbc07f-e1d9-452f-b859-a0d647a0a8a2 · outbound

This paper cites There is no SAMantics! Exploring SAM as a Backbone for Visual Understanding Tasks.

AVA-Bench: Atomic Visual Ability Benchmark for Vision Foundation Models There is no SAMantics! Exploring SAM as a Backbone for Visual Understanding Tasks

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-19T11:13:02.883068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T11:12:41.130806Z digest=sha256:77dcab7fee78d864f54a598b0f22ccc964ceeb76b01e817e0177d5f09dacd1d0

Observation 14ee273e-135e-47aa-9a9f-7a009c3d6e88 · outbound

This paper cites MLLM-SUL: Multimodal Large Language Model for Semantic Scene Understanding and Localization in Traffic Scenarios.

AVA-Bench: Atomic Visual Ability Benchmark for Vision Foundation Models MLLM-SUL: Multimodal Large Language Model for Semantic Scene Understanding and Localization in Traffic Scenarios

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-19T11:13:02.893274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T11:12:41.130806Z digest=sha256:4c808fbcd6a935c2e52fcf730102863bab24ffa5f1ff19bd543e690ea264f296

Observation 8f3a1ec8-383f-4a72-92bd-bacd9751eb5e · outbound

This paper cites Multimodal Autoregressive Pre-training of Large Vision Encoders.

AVA-Bench: Atomic Visual Ability Benchmark for Vision Foundation Models Multimodal Autoregressive Pre-training of Large Vision Encoders

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-19T11:13:02.851822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T11:12:41.130806Z digest=sha256:fcd9bcffbfb2b8893cf4425c67037b70956beef19b8695045003d9659dfb6f5f

Observation c7e1e9ca-418f-4a91-ae97-22de8b9f90ac · outbound

This paper cites Blink: Multimodal large language models can see but not perceive.

AVA-Bench: Atomic Visual Ability Benchmark for Vision Foundation Models Blink: Multimodal large language models can see but not perceive

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T11:13:03.089464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T11:12:41.130806Z digest=sha256:96411834d34affd96bf39900c7b6e713a505fdc3906dd99933e1b2a2366a9510

Observation 6b402099-349e-4aec-86e3-ab2dafd7eae3 · outbound

This paper cites Can We Talk Models Into Seeing the World Differently?.

AVA-Bench: Atomic Visual Ability Benchmark for Vision Foundation Models Can We Talk Models Into Seeing the World Differently?

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-19T11:13:02.855394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T11:12:41.130806Z digest=sha256:655a56959b64a6939420fdd9d96ad72b4a68a4054e249e7c1fefee6f537c4908

Observation 1349b0e5-2b4d-4536-acce-60f3a83c11c7 · outbound

This paper cites Vision meets robotics: The kitti dataset.

AVA-Bench: Atomic Visual Ability Benchmark for Vision Foundation Models Vision meets robotics: The kitti dataset

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T11:13:03.049252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T11:12:41.130806Z digest=sha256:7ca92cf05e9712b650782c887546a43483a9cf2b8f7d5cec6b1ec80399754d55

Observation 84c79353-0db4-4b05-8878-4e830072408b · outbound

This paper cites Battle of the backbones: A large- scale comparison of pretrained models across computer vision tasks.

AVA-Bench: Atomic Visual Ability Benchmark for Vision Foundation Models Battle of the backbones: A large- scale comparison of pretrained models across computer vision tasks

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T11:13:03.098323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T11:12:41.130806Z digest=sha256:570fff2f46143c352d9145b4a223c0e64d5d6fed0d9613790f9733db8c13ce06

Observation 71a977c4-dd80-4fb7-a8ec-e2e7c1c92d02 · outbound

This paper cites Making the v in vqa matter: Elevating the role of image understanding in visual question answering.

AVA-Bench: Atomic Visual Ability Benchmark for Vision Foundation Models Making the v in vqa matter: Elevating the role of image understanding in visual question answering

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T11:13:03.113924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T11:12:41.130806Z digest=sha256:daefcb556b314b58e25baae6f8bdf154bb20ef0afd3ef6c39c2667160600084e

Observation dab1989f-8231-43f9-901a-70790b83010a · outbound

This paper cites Making the V in VQA matter: Elevating the role of image understanding in Visual Question Answering.

AVA-Bench: Atomic Visual Ability Benchmark for Vision Foundation Models Making the V in VQA matter: Elevating the role of image understanding in Visual Question Answering

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T11:13:03.059969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T11:12:41.130806Z digest=sha256:4c0d20c0e722835b8ecd76805bacb34f5b0c61f532785262fa980a9f05bf9541

Observation 761f0e06-25f8-4d87-8ab4-5b96ba627dcf · outbound

This paper cites The Llama 3 Herd of Models.

AVA-Bench: Atomic Visual Ability Benchmark for Vision Foundation Models The Llama 3 Herd of Models

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-05-19T11:13:02.808841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T11:12:41.130806Z digest=sha256:e53f83127bfd759d3e099e8a2b850f59e6f1994833c78b5f8f995342c7481d17

Observation a23c15d7-6a06-4f51-919b-64c884ff4908 · outbound

This paper cites Bioclip 2: Emergent properties from scaling hierarchical contrastive learning.

AVA-Bench: Atomic Visual Ability Benchmark for Vision Foundation Models Bioclip 2: Emergent properties from scaling hierarchical contrastive learning

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-19T11:13:02.798927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T11:12:41.130806Z digest=sha256:4b7b0ab6dfaab263d16e1c1bdc3cf6255bd34d169982bcfd953e2bdee0b1a7df

Observation 938290bf-92e8-4df5-b001-a3d7ad78e1af · outbound

This paper cites Lvis: A dataset for large vocabulary instance segmentation.

AVA-Bench: Atomic Visual Ability Benchmark for Vision Foundation Models Lvis: A dataset for large vocabulary instance segmentation

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T11:13:03.047900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T11:12:41.130806Z digest=sha256:a4c5cce7125a7bfa84073a8d8fadd23b2d7a4cafe6e9bb0eac652206acc128c9

Observation 10c56288-30be-4fef-95c6-3fbda059ac18 · outbound

This paper cites A survey on vision transformer.

AVA-Bench: Atomic Visual Ability Benchmark for Vision Foundation Models A survey on vision transformer

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T11:13:03.036431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T11:12:41.130806Z digest=sha256:09687add2dd478b63a4de158081164be1bde27183a573acd67662739141af7ec

Observation 581289eb-d9d6-4aba-9543-f16e2f48ab58 · outbound

This paper cites RADIOv2.5: Improved Baselines for Agglomerative Vision Foundation Models.

AVA-Bench: Atomic Visual Ability Benchmark for Vision Foundation Models RADIOv2.5: Improved Baselines for Agglomerative Vision Foundation Models

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-19T11:13:02.805685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T11:12:41.130806Z digest=sha256:21858b37f2a493893d3d26a6ab0608b1be6c64c19150dfe19a30bbabe6ed4fdc

Observation 587cdc6e-f774-46de-b132-eaef90308a44 · outbound

This paper cites Parameter-efficient transfer learning for nlp.

AVA-Bench: Atomic Visual Ability Benchmark for Vision Foundation Models Parameter-efficient transfer learning for nlp

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T11:13:03.014019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T11:12:41.130806Z digest=sha256:e56a9ffe9bf9c3dacb4abc1f867ff8a8cf2ea54bbfb329a5d5dcef6fc2dfae4a

Observation 12c3d566-fafb-456e-9f18-c08253b40c1e · outbound

This paper cites Drone-based object counting by spatially regularized regional proposal network.

AVA-Bench: Atomic Visual Ability Benchmark for Vision Foundation Models Drone-based object counting by spatially regularized regional proposal network

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T11:13:03.021238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T11:12:41.130806Z digest=sha256:e3f4e8305870c4a8829d1f88a5f73561a6c24039cbcf39735f496fc45a0bb065

Observation ec97031c-663a-4160-872b-faab2438c3e8 · outbound

This paper cites Lora: Low-rank adaptation of large language models.

AVA-Bench: Atomic Visual Ability Benchmark for Vision Foundation Models Lora: Low-rank adaptation of large language models

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T11:13:03.113477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T11:12:41.130806Z digest=sha256:f24cfc078f38a3c0c7c7a231d1eec1ac78161084f1321c118ca6547a8413c042

Observation f09d3e57-92b7-4da0-9496-38c4d707aeac · outbound

This paper cites A Survey on Evaluation of Multimodal Large Language Models.

AVA-Bench: Atomic Visual Ability Benchmark for Vision Foundation Models A Survey on Evaluation of Multimodal Large Language Models

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-19T11:13:02.831115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T11:12:41.130806Z digest=sha256:3e324b985c974996ed7ba6e92efe1744611c46682ccec49e15b4889959af4349

Observation 2f3652a3-a196-47f5-8dfe-c0797ad7d866 · outbound

This paper cites T2i-compbench: A comprehensive benchmark for open-world compositional text-to-image generation.

AVA-Bench: Atomic Visual Ability Benchmark for Vision Foundation Models T2i-compbench: A comprehensive benchmark for open-world compositional text-to-image generation

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T11:13:03.079154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T11:12:41.130806Z digest=sha256:8a2ef1c1455a6567795feac65ce5c8a5134d65e2f284cb312980b5c55aa698f2

Observation 3b18747c-e5c2-4e08-b29d-f0194bff28b1 · outbound

This paper cites OCR-Reasoning Benchmark: Unveiling the True Capabilities of MLLMs in Complex Text-Rich Image Reasoning.

AVA-Bench: Atomic Visual Ability Benchmark for Vision Foundation Models OCR-Reasoning Benchmark: Unveiling the True Capabilities of MLLMs in Complex Text-Rich Image Reasoning

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-27T02:04:28.324233Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T11:12:41.130806Z digest=sha256:49fa4ef8eacee81cd7b662efdc5ed1887c6016f73eb4b03a67a43e3caf8224ad

Observation 0adf2cfe-40d1-4484-a6cd-278204a07da3 · outbound

This paper cites Position: The platonic representation hypothesis.

AVA-Bench: Atomic Visual Ability Benchmark for Vision Foundation Models Position: The platonic representation hypothesis

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T11:13:03.105719Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T11:12:41.130806Z digest=sha256:15c144737ac1cf60d01eed9811ff8b784ca294314b04637455e0f8c075ff242f

Observation f9d7873b-8f6f-47f2-8402-09bbd1246c7f · outbound

This paper cites Is 'Right' Right? Enhancing Object Orientation Understanding in Multimodal Large Language Models through Egocentric Instruction Tuning.

AVA-Bench: Atomic Visual Ability Benchmark for Vision Foundation Models Is 'Right' Right? Enhancing Object Orientation Understanding in Multimodal Large Language Models through Egocentric Instruction Tuning

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-19T11:13:02.890631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T11:12:41.130806Z digest=sha256:075bf2191ca4718977bc51afeb2b591ff142822ef31e7acac5588c50f8848cc8

Observation c088ac66-75f3-42bb-8ec2-5d2aeabd789f · outbound

This paper cites Transformers in vision: A survey.

AVA-Bench: Atomic Visual Ability Benchmark for Vision Foundation Models Transformers in vision: A survey

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T11:13:03.056263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T11:12:41.130806Z digest=sha256:0892dfb2bd19e2ad7e0b8f8a8291017ad9ec4afafe17752cd229a8a251358b48

Observation 90af026b-3dbf-4772-bc01-fbf76c449e15 · outbound

This paper cites Mllm-compbench: A comparative reasoning benchmark for multimodal llms.

AVA-Bench: Atomic Visual Ability Benchmark for Vision Foundation Models Mllm-compbench: A comparative reasoning benchmark for multimodal llms

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T11:13:03.027988Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T11:12:41.130806Z digest=sha256:550d85d70b029f097774d671986b6937323abcef5a637a4ba51794e3f0717e78

Observation d60c2d93-3d20-4f62-b0e5-6a5ccfc3a5f6 · outbound

This paper cites Finer: Investigating and enhancing fine-grained visual concept recognition in large vision language models.

AVA-Bench: Atomic Visual Ability Benchmark for Vision Foundation Models Finer: Investigating and enhancing fine-grained visual concept recognition in large vision language models

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T11:13:03.100148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T11:12:41.130806Z digest=sha256:29b36de4902a87af8d4345fea6af58da8865b19b90b4171a9daefd242f934695

Observation c4205e3f-d4cf-4898-a504-05a1c8e043fc · outbound

This paper cites Segment anything.

AVA-Bench: Atomic Visual Ability Benchmark for Vision Foundation Models Segment anything

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T11:13:03.063129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T11:12:41.130806Z digest=sha256:58935a0f46d825d1087d80ee4559897f824343d5d370e2dd4ce577785fce8616

Observation 2780cd61-b304-4fe8-b521-a2eb3f05ae84 · outbound

This paper cites Diagnostics- llava: A visual language model for domain-specific diagnostics of equipment.

AVA-Bench: Atomic Visual Ability Benchmark for Vision Foundation Models Diagnostics- llava: A visual language model for domain-specific diagnostics of equipment

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T11:13:03.112116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T11:12:41.130806Z digest=sha256:dd6c2bf724564324735a375ff81a2725d889133f7ee96c1680ae0d8f01512429

Observation 36929f7f-0e8c-4f36-93ac-32430d2e1063 · outbound

This paper cites Kylberg texture dataset v.

AVA-Bench: Atomic Visual Ability Benchmark for Vision Foundation Models Kylberg texture dataset v

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T11:13:03.073553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T11:12:41.130806Z digest=sha256:f20c4852bca5bcc0145b424fdc50d7ad9a5d66977e09f616d0156b2610881c3d

Observation 91724a07-6f19-46b8-862a-18de7a7caa48 · outbound

This paper cites LLMCount: Enhancing Stationary mmWave Detection with Multimodal-LLM.

AVA-Bench: Atomic Visual Ability Benchmark for Vision Foundation Models LLMCount: Enhancing Stationary mmWave Detection with Multimodal-LLM

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-19T11:13:02.815623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T11:12:41.130806Z digest=sha256:c14d9d07e90d3d522bfca1e4e138b6096986a04e17e391c70728cb12de664302

Observation 2f785da9-4afd-42d3-9533-0324a7c94836 · outbound

This paper cites EALD-MLLM: Emotion Analysis in Long-sequential and De-identity videos with Multi-modal Large Language Model.

AVA-Bench: Atomic Visual Ability Benchmark for Vision Foundation Models EALD-MLLM: Emotion Analysis in Long-sequential and De-identity videos with Multi-modal Large Language Model

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-19T11:13:02.783885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T11:12:41.130806Z digest=sha256:a0b68a7b979227cd8f29eaef22b66b8c6619bca5a13877040e08d0f0d8fe5a6b

Observation 4d6fc00e-e34a-432f-ad90-da15443ca303 · outbound

This paper cites Video crowd localization with multifocus gaussian neighborhood attention and a large-scale benchmark.

AVA-Bench: Atomic Visual Ability Benchmark for Vision Foundation Models Video crowd localization with multifocus gaussian neighborhood attention and a large-scale benchmark

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T11:13:03.034870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T11:12:41.130806Z digest=sha256:dd5264514663d30d3405e9d7db9b4568476d043a919b33edb29871d33ed804fa

Observation 92e62d28-3d9d-4cbc-816d-387002fdbab4 · outbound

This paper cites Object detection in optical remote sensing images: A survey and a new benchmark.

AVA-Bench: Atomic Visual Ability Benchmark for Vision Foundation Models Object detection in optical remote sensing images: A survey and a new benchmark

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T11:13:03.024570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T11:12:41.130806Z digest=sha256:7843299ed8aad3506db052ad487c03915910f34ec32044e4cfaa9f3cf602ee07

Observation bee65464-8792-4404-b280-dd0755252bec · outbound

This paper cites Reliable crowdsourcing and deep locality-preserving learning for un- constrained facial expression recognition.

AVA-Bench: Atomic Visual Ability Benchmark for Vision Foundation Models Reliable crowdsourcing and deep locality-preserving learning for un- constrained facial expression recognition

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T11:13:03.084431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T11:12:41.130806Z digest=sha256:2d9cbb3350547b83b8763895a4dace956ae7c573e9f18d5ca0d088cf7f536118

Observation d4b948b0-1000-47c2-baf9-0a7d7665ae36 · outbound

This paper cites OpenVision: A Fully-Open, Cost-Effective Family of Advanced Vision Encoders for Multimodal Learning.

AVA-Bench: Atomic Visual Ability Benchmark for Vision Foundation Models OpenVision: A Fully-Open, Cost-Effective Family of Advanced Vision Encoders for Multimodal Learning

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-19T11:13:02.865609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T11:12:41.130806Z digest=sha256:f124a01786994656fc36877062cfa121d453d130578e8ebe53cc27cd9caf316a

Observation 78c8ddaa-47de-4843-8a1b-958b2fea1532 · outbound

This paper cites Visual Large Language Models for Generalized and Specialized Applications.

AVA-Bench: Atomic Visual Ability Benchmark for Vision Foundation Models Visual Large Language Models for Generalized and Specialized Applications

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-19T11:13:02.863394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T11:12:41.130806Z digest=sha256:aef50521c5ce6dc22e6ae26ade6067690a07d3e25f21f1d34a36dc1b9e3c5334

Observation 235fe6d8-1aa0-4bd6-a9f3-78d964c14141 · outbound

This paper cites Expression analysis based on face regions in real-world conditions.

AVA-Bench: Atomic Visual Ability Benchmark for Vision Foundation Models Expression analysis based on face regions in real-world conditions

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T11:13:03.096485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T11:12:41.130806Z digest=sha256:a2c642eff8e127487ed68f51a87034371686971dc362579b9f88060844559e8a

Observation 2c457a82-0bc7-4ae8-a8da-7d3184341a90 · outbound

This paper cites A comprehensive survey and guide to multimodal large language models in vision-language tasks.

AVA-Bench: Atomic Visual Ability Benchmark for Vision Foundation Models A comprehensive survey and guide to multimodal large language models in vision-language tasks

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-19T11:13:02.833937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T11:12:41.130806Z digest=sha256:2a7b2dac1d5def53d5b0d12683d7e22160b0cde08b2963549a69bff94e453aea

Observation 5634bf49-ca72-46da-807e-9fb3bc46e823 · outbound

This paper cites Visual instruction tuning.

AVA-Bench: Atomic Visual Ability Benchmark for Vision Foundation Models Visual instruction tuning

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T11:13:03.109478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T11:12:41.130806Z digest=sha256:7f8192ad7795ecce1c5ec2f5c2baf17fd1b19d9b358cc4605d73c4bd25f410f4

Observation eaa0e79a-a998-4364-b3c5-60d2f61f302b · outbound

This paper cites Ocrbench: on the hidden mystery of ocr in large multimodal models.

AVA-Bench: Atomic Visual Ability Benchmark for Vision Foundation Models Ocrbench: on the hidden mystery of ocr in large multimodal models

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T11:13:03.012035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T11:12:41.130806Z digest=sha256:a822a9bae56e24b2d02373efdfddd1ede01838ccd0e635d645fc68771b195751

Observation 629a158d-0e98-4440-aa15-752af39e5a47 · outbound

This paper cites Cvpr 2020 continual learning in computer vision competition: Approaches, results, current challenges and future directions.

AVA-Bench: Atomic Visual Ability Benchmark for Vision Foundation Models Cvpr 2020 continual learning in computer vision competition: Approaches, results, current challenges and future directions

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T11:13:03.023017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T11:12:41.130806Z digest=sha256:5aa2ae19ddaee0a54de11d55c0fbc0fd93cbacc262d3229dcb533ce4a1721d74

Observation 7959ce49-47a4-46f1-92a9-9abd5f30d7b4 · outbound

This paper cites The development of the cie 2000 colour-difference formula: Ciede2000.

AVA-Bench: Atomic Visual Ability Benchmark for Vision Foundation Models The development of the cie 2000 colour-difference formula: Ciede2000

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T11:13:03.115407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T11:12:41.130806Z digest=sha256:a996f7cfba7f99f70b1eb759d29cc90dcd4c261f9fcefae082fe3da512d46976

Observation 9d991c58-b310-4351-8e5e-b96c42ec8568 · outbound

This paper cites Fine-tuning is fine, if calibrated.

AVA-Bench: Atomic Visual Ability Benchmark for Vision Foundation Models Fine-tuning is fine, if calibrated

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T11:13:03.122446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T11:12:41.130806Z digest=sha256:6ef1aa67d937a55e3b87458ee825d67dd3d5017da2429e3b8266bc270a81f77e

Observation 90165c19-b3b1-419c-a4fb-96263d7bc45a · outbound

This paper cites Batch-level Experience Replay with Review for Continual Learning.

AVA-Bench: Atomic Visual Ability Benchmark for Vision Foundation Models Batch-level Experience Replay with Review for Continual Learning

Reference 61

Resolution
verified exact
arxiv_id, observed 2026-05-19T11:13:02.840356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T11:12:41.130806Z digest=sha256:84b9a8a176f49f0b88d3c0c51456ddbe2ba8f9b15e311f95c3076bb6ea8eb91a

Observation bc806bbd-a42e-45c5-8bd5-0929314470f8 · outbound

This paper cites Online continual learning in image classification: An empirical survey.

AVA-Bench: Atomic Visual Ability Benchmark for Vision Foundation Models Online continual learning in image classification: An empirical survey

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T11:13:03.069356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T11:12:41.130806Z digest=sha256:074bcb598153f2f3214072b3a6147edc23a3936f94abace27711d4d2a2a4a823

Observation 63a3c159-cb68-4ac3-9f4d-fd6dc5045147 · outbound

This paper cites Supervised contrastive replay: Revisiting the nearest class mean classifier in online class-incremental continual learning.

AVA-Bench: Atomic Visual Ability Benchmark for Vision Foundation Models Supervised contrastive replay: Revisiting the nearest class mean classifier in online class-incremental continual learning

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T11:13:03.107677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T11:12:41.130806Z digest=sha256:2ba9e9f0ce18b15ede86f8be1eb0931dc92de4d71e118cc5b85baf094d9e4174

Observation 258241d7-3370-43ee-a30f-ed82d10371bc · outbound

This paper cites Lessons and insights from a unifying study of parameter-efficient fine-tuning (peft) in visual recognition.

AVA-Bench: Atomic Visual Ability Benchmark for Vision Foundation Models Lessons and insights from a unifying study of parameter-efficient fine-tuning (peft) in visual recognition

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T11:13:03.104107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T11:12:41.130806Z digest=sha256:492fdd3f8f78cc1cfdb306488d4da9730258bd824f872882ba057d0b120cfdb9

Observation 3e61a3dc-c11f-49cd-b0f1-a526db319c9c · outbound

This paper cites Fine-Grained Visual Classification of Aircraft.

AVA-Bench: Atomic Visual Ability Benchmark for Vision Foundation Models Fine-Grained Visual Classification of Aircraft

Reference 65

Resolution
verified exact
local_arxiv, observed 2026-05-19T11:13:02.885453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T11:12:41.130806Z digest=sha256:76c0a6f55454e0821bf8556571170d445b14a419b844b3ad9cd2e4ee624c21bb

Observation 58126c4a-9a57-4fff-acee-edffef7b7dbe · outbound

This paper cites The kth-tips2 database.

AVA-Bench: Atomic Visual Ability Benchmark for Vision Foundation Models The kth-tips2 database

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T11:13:03.074240Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T11:12:41.130806Z digest=sha256:003c53373286e60795c9f358c39a79acecdb86112e2df394e30e9a9e5d438af7

Observation c52d5607-8014-40b7-9097-c57445bb0ab4 · outbound

This paper cites Hierarchical interpretable vision reasoning driven through a multi-modal large language model for depth estimation.

AVA-Bench: Atomic Visual Ability Benchmark for Vision Foundation Models Hierarchical interpretable vision reasoning driven through a multi-modal large language model for depth estimation

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T11:13:03.071893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T11:12:41.130806Z digest=sha256:7244524f212172ef850b2d751db62aa90b4f0cede2c9ecb96c87f12704e39d52

Observation c47b2d68-de1e-4a29-9b69-d2fcb5abfb89 · outbound

This paper cites Mishra, K.

AVA-Bench: Atomic Visual Ability Benchmark for Vision Foundation Models Mishra, K

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T11:13:03.070887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T11:12:41.130806Z digest=sha256:94514f78fb043141431a37c952444dcf8cd6c3b64dbdb37a363dd8fe00eebe5d

Observation 81a9079f-32f9-43b1-9692-5c1966ee7e75 · outbound

This paper cites Moments in time dataset: one million videos for event understanding.

AVA-Bench: Atomic Visual Ability Benchmark for Vision Foundation Models Moments in time dataset: one million videos for event understanding

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T11:13:03.087860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T11:12:41.130806Z digest=sha256:dc0ebd8c86ef3f28d4aef0ccfecf131387b485b80736ced0da8cd69a424ee17a

Observation f407c0bc-cd9f-4746-b490-ae7dfeeaf049 · outbound

This paper cites Intriguing properties of vision transformers.

AVA-Bench: Atomic Visual Ability Benchmark for Vision Foundation Models Intriguing properties of vision transformers

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T11:13:03.091320Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T11:12:41.130806Z digest=sha256:f38902a9143bdbd8fe028cb21467fa9ea69d2b6f88c9a693636aa8cedbc32a7e

Observation 2bda31ff-adde-4223-ae84-e3c5287a7ff7 · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

AVA-Bench: Atomic Visual Ability Benchmark for Vision Foundation Models DINOv2: Learning Robust Visual Features without Supervision

Reference 71

Resolution
verified exact
local_arxiv, observed 2026-05-19T11:13:02.898494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T11:12:41.130806Z digest=sha256:f55f86e824094d1fa5f07548fc849943a74c0634892aa72b1601381a98842b09

Observation 2d451229-ee15-4655-8d1c-ddbf50c64bb7 · outbound

This paper cites A simple interpretable transformer for fine-grained image classification and analysis.

AVA-Bench: Atomic Visual Ability Benchmark for Vision Foundation Models A simple interpretable transformer for fine-grained image classification and analysis

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T11:13:03.091054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T11:12:41.130806Z digest=sha256:dfa75417c57229811d2133b2f23ff0526648de112b2d0e763e3790bed3556c90

Observation 288c74a6-29f8-467e-bbfc-8561dc9f974f · outbound

This paper cites Learning transferable visual models from natural language supervision.

AVA-Bench: Atomic Visual Ability Benchmark for Vision Foundation Models Learning transferable visual models from natural language supervision

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T11:13:03.031518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T11:12:41.130806Z digest=sha256:13ed2e58d330ea53ce7f654b6ef9a1ba3f6797c963b01529eddbda018b7bdacc

Observation c5a373b1-63b1-49f8-8f15-9c4fe9ceb839 · outbound

This paper cites Towards robust monocular depth estimation: Mixing datasets for zero-shot cross-dataset transfer.

AVA-Bench: Atomic Visual Ability Benchmark for Vision Foundation Models Towards robust monocular depth estimation: Mixing datasets for zero-shot cross-dataset transfer

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T11:13:03.045654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T11:12:41.130806Z digest=sha256:59e576350d5f52b618c6435863dada455742ccc7ba1c2081108416d2d87a96d5

Observation 61e4a709-cd34-4325-abbb-3a485ace0b03 · outbound

This paper cites Learning to count everything.

AVA-Bench: Atomic Visual Ability Benchmark for Vision Foundation Models Learning to count everything

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T11:13:03.070155Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T11:12:41.130806Z digest=sha256:8d36ff4939165f9aa3caa61af3935cc06234b4e2c4e4ec6ad019e658fdcfae16

Observation ee6569a7-5989-4dbb-8918-d7ea671eb1f5 · outbound

This paper cites Am-radio: Agglomerative vision foundation model reduce all domains into one.

AVA-Bench: Atomic Visual Ability Benchmark for Vision Foundation Models Am-radio: Agglomerative vision foundation model reduce all domains into one

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T11:13:03.094691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T11:12:41.130806Z digest=sha256:cb3ae5fd0d26fd5fefdca4a84c3853ec1f1c2a7a0b2274a83c9116196e96f861

Observation 0eb6ed2e-cad1-428b-ba80-9e926668f3ba · outbound

This paper cites Generalized intersection over union: A metric and a loss for bounding box regression.

AVA-Bench: Atomic Visual Ability Benchmark for Vision Foundation Models Generalized intersection over union: A metric and a loss for bounding box regression

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T11:13:03.054686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T11:12:41.130806Z digest=sha256:5725521ec7c651ea8119a762358b4d7ddacc49d30faaef722b3cc88bd8875813

Observation 529d4ca2-5bc1-4922-8da1-c1ca9a9287dd · outbound

This paper cites Object detection with multimodal large vision-language models: An in-depth review.

AVA-Bench: Atomic Visual Ability Benchmark for Vision Foundation Models Object detection with multimodal large vision-language models: An in-depth review

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T11:13:03.033473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T11:12:41.130806Z digest=sha256:38224cf1e7160db886c4e5500bd97eed10eb19309c5c45b0df297e63c43f7284

Observation 0e2a7e22-46de-41bf-8050-c79d6cd5a24d · outbound

This paper cites Objects365: A large-scale, high-quality dataset for object detection.

AVA-Bench: Atomic Visual Ability Benchmark for Vision Foundation Models Objects365: A large-scale, high-quality dataset for object detection

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T11:13:03.022768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T11:12:41.130806Z digest=sha256:94e605cbfbe5a403cd3e00102f965d6d5df9ef24008e47654f029e260620e7d6

Observation 5fc250af-a27a-4e9c-aa01-363cd7d71db9 · outbound

This paper cites Online class- incremental continual learning with adversarial shapley value.

AVA-Bench: Atomic Visual Ability Benchmark for Vision Foundation Models Online class- incremental continual learning with adversarial shapley value

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T11:13:03.100342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T11:12:41.130806Z digest=sha256:14877f7df3cbe3863c6f34aad44e8d9f7c54c8d476311e5adba709a2f2619ecf

Observation a837bc62-ca33-4516-92ff-35b87f2b68ec · outbound

This paper cites Indoor segmentation and support inference from rgbd images.

AVA-Bench: Atomic Visual Ability Benchmark for Vision Foundation Models Indoor segmentation and support inference from rgbd images

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T11:13:03.001857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T11:12:41.130806Z digest=sha256:a1082c26de4bcdaadd3d356763b5bf93937f72a71d5c2ec0d1775fec765b2505

Observation bac7d0bb-d7cf-4d9a-91cd-326c1898d884 · outbound

This paper cites Multi-modal large language models are effective vision learners.

AVA-Bench: Atomic Visual Ability Benchmark for Vision Foundation Models Multi-modal large language models are effective vision learners

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T11:13:03.045962Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T11:12:41.130806Z digest=sha256:759d6e1254944bae7b10c37f3ea3b54f4c48d1e12bedb02752a64614003659de

Observation a4c18339-ab8b-43aa-830f-6c9bcba4d52a · outbound

This paper cites Grounding multimodal large language models in actions.

AVA-Bench: Atomic Visual Ability Benchmark for Vision Foundation Models Grounding multimodal large language models in actions

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T11:13:03.119018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T11:12:41.130806Z digest=sha256:b95d3c0bd698c87e0314f11df835d92821b6c770bdf96506903ed76192543c4a

Observation d36839aa-c845-46ea-a942-ba82fe55cdfc · outbound

This paper cites Temel, J.

AVA-Bench: Atomic Visual Ability Benchmark for Vision Foundation Models Temel, J

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T11:13:03.041863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T11:12:41.130806Z digest=sha256:ce1f1ae13c6a89e082ddda3c4cbeb3683dc660e234463327a1b0cd0c106f1a1e

Observation f88cff89-3b75-46f4-8f65-0845fa473fc7 · outbound

This paper cites Kth-tips dataset.

AVA-Bench: Atomic Visual Ability Benchmark for Vision Foundation Models Kth-tips dataset

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T11:13:03.042244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T11:12:41.130806Z digest=sha256:21c0667bee2713020839d0882fe2dac74c7bc17a83a9847f128e04839a79c7ee

Observation dc319ec0-c93b-4a95-9bd0-8a77737afcc8 · outbound

This paper cites Semantic segmentation using vision transformers: A survey.

AVA-Bench: Atomic Visual Ability Benchmark for Vision Foundation Models Semantic segmentation using vision transformers: A survey

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T11:13:03.026231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T11:12:41.130806Z digest=sha256:62a99d30981b0386e24cd6dbc301ef7db750ea3624cca15565fb7c45504041e7

Observation df647130-248f-4f48-8cac-d6ddab8075af · outbound

This paper cites Cambrian-1: A fully open, vision-centric exploration of multimodal llms.

AVA-Bench: Atomic Visual Ability Benchmark for Vision Foundation Models Cambrian-1: A fully open, vision-centric exploration of multimodal llms

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T11:13:03.111846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T11:12:41.130806Z digest=sha256:b0d870afe9f38c60e852a929b1159766fb23e7bcb622708f7aaec5b6727bc22d

Observation ac95146e-21d8-4074-b945-b0c5c5d71d14 · outbound

This paper cites Eyes wide shut? exploring the visual shortcomings of multimodal llms.

AVA-Bench: Atomic Visual Ability Benchmark for Vision Foundation Models Eyes wide shut? exploring the visual shortcomings of multimodal llms

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T11:13:02.991510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T11:12:41.130806Z digest=sha256:6486bbc2894fc315f11f286e7411f25c45080fdec3a492b30a0f1592bf0a4216

Observation d3ed9224-d83b-4860-9309-d5ac4fec49e8 · outbound

This paper cites SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features.

AVA-Bench: Atomic Visual Ability Benchmark for Vision Foundation Models SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features

Reference 89

Resolution
verified exact
local_arxiv, observed 2026-05-19T11:13:02.844036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T11:12:41.130806Z digest=sha256:0c6156cfd6765df8f1478962a67c0e571080091cd581a8f3d34b0570c598248f

Observation 334a8d07-268e-4315-b5cc-672a93d8d441 · outbound

This paper cites Holistic transfer: Towards non-disruptive fine-tuning with partial target data.

AVA-Bench: Atomic Visual Ability Benchmark for Vision Foundation Models Holistic transfer: Towards non-disruptive fine-tuning with partial target data

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T11:13:03.101869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T11:12:41.130806Z digest=sha256:3ab0dde563ea1064510f192b4e6698a6fe32b01f504d77f6122dce4a640af72c

Observation 96ca5726-1e9e-46b6-a194-c8a2ec7c782b · outbound

This paper cites Visual query tuning: Towards effective usage of intermediate representations for parameter and memory efficient transfer learning.

AVA-Bench: Atomic Visual Ability Benchmark for Vision Foundation Models Visual query tuning: Towards effective usage of intermediate representations for parameter and memory efficient transfer learning

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T11:13:03.051304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T11:12:41.130806Z digest=sha256:6566a002d3c6a2b7ca8b4eef686bddfcfc61aea1c0e3cf323ae22e8b3b8b2397

Observation e962a679-52ab-4177-a821-94f2673b8beb · outbound

This paper cites Benchmarking representation learning for natural world image collections.

AVA-Bench: Atomic Visual Ability Benchmark for Vision Foundation Models Benchmarking representation learning for natural world image collections

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T11:13:03.092775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T11:12:41.130806Z digest=sha256:e2675826eba54452c57a638199180c17112bfe22c27b53ae2c6c55a246721f20

Observation 2c669783-fe6d-4e56-ba56-abe50041330b · outbound

This paper cites COCO-Text: Dataset and Benchmark for Text Detection and Recognition in Natural Images.

AVA-Bench: Atomic Visual Ability Benchmark for Vision Foundation Models COCO-Text: Dataset and Benchmark for Text Detection and Recognition in Natural Images

Reference 93

Resolution
verified exact
local_arxiv, observed 2026-05-19T11:13:02.871922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T11:12:41.130806Z digest=sha256:29e7fd96d3cbd7a396b4b2d91a3cde161976056680c74ca9fc4f20e6edc77aa6

Observation 12569ffd-69b6-444f-a391-b50780b989e7 · outbound

This paper cites The caltech-ucsd birds-200-2011 dataset.

AVA-Bench: Atomic Visual Ability Benchmark for Vision Foundation Models The caltech-ucsd birds-200-2011 dataset

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T11:13:03.054934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T11:12:41.130806Z digest=sha256:58ec929562f1afa60b6c795e912c729695a236f6383326bf70d74275bf5b7c39

Observation 28a70ba5-ed18-4be6-b4e3-910923c938c0 · outbound

This paper cites Harnessing multi-modal large language models for measuring and interpreting color differences.

AVA-Bench: Atomic Visual Ability Benchmark for Vision Foundation Models Harnessing multi-modal large language models for measuring and interpreting color differences

Reference 95

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T11:13:03.096667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T11:12:41.130806Z digest=sha256:c447e8803571efa6c20433121e0c643a49e69836564bba8b8640510d5b475383

Observation 3cad97b3-0415-4738-a46f-b2c110d9613b · outbound

This paper cites Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence.

AVA-Bench: Atomic Visual Ability Benchmark for Vision Foundation Models Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence

Reference 96

Resolution
verified exact
local_arxiv, observed 2026-05-19T11:13:02.859921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T11:12:41.130806Z digest=sha256:77f7b0d2681c151232736d4e80c6b1d223e5c86d4aa7a97f32c54a397335d19a

Observation d5c2b3b6-725a-4bcb-b5d1-1a95e099416d · outbound

This paper cites Visionllm v2: An end-to-end generalist multimodal large language model for hundreds of vision-language tasks.

AVA-Bench: Atomic Visual Ability Benchmark for Vision Foundation Models Visionllm v2: An end-to-end generalist multimodal large language model for hundreds of vision-language tasks

Reference 97

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T11:13:03.076991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T11:12:41.130806Z digest=sha256:0695e45e2a8dd1674f2b8aa68c6c7852482959bd22e615142f92d7ca17ee034a

Observation ff3c0eb4-d383-4fb5-86d5-06f5b357602a · outbound

This paper cites Multimodal large language models: A survey.

AVA-Bench: Atomic Visual Ability Benchmark for Vision Foundation Models Multimodal large language models: A survey

Reference 98

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T11:13:03.094882Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T11:12:41.130806Z digest=sha256:1203d6b8a797bd8967153d3aa602502f4d41299d35db6cf22820d59995e6a6b6

Observation 8a9d4413-ca5d-4590-968c-e60a853897f6 · outbound

This paper cites Visual Compositional Tuning.

AVA-Bench: Atomic Visual Ability Benchmark for Vision Foundation Models Visual Compositional Tuning

Reference 99

Resolution
verified exact
arxiv_id, observed 2026-05-20T00:00:19.480510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T11:12:41.130806Z digest=sha256:8f2e5e0a4d1a576a5fe668eb992c9e906fb9443c3d58d3d5cbee5a4a6400d98d

Observation 61dff8f3-c356-46eb-9b8e-a662c59e2651 · outbound

This paper cites ConceptMix: A Compositional Image Generation Benchmark with Controllable Difficulty.

AVA-Bench: Atomic Visual Ability Benchmark for Vision Foundation Models ConceptMix: A Compositional Image Generation Benchmark with Controllable Difficulty

Reference 100

Resolution
verified exact
arxiv_id, observed 2026-05-19T11:13:02.865956Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T11:12:41.130806Z digest=sha256:695f9e5b76f8f827f34d56c376f412f39f97e8164b3c93f9313326624d5f2c21

Pith citing papers

Observation 5e2c7784-6a4f-44c5-9b7c-acb1a08a2ebb · inbound

Revisiting Model Stitching In the Foundation Model Era cites this paper.

Revisiting Model Stitching In the Foundation Model Era AVA-Bench: Atomic Visual Ability Benchmark for Vision Foundation Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-07-14T22:19:38.973091Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T22:19:38.973091Z digest=sha256:2edbc4a898559b8affe97e3b1a16548770be5e9da0aa59dde3affa76b6d4d9bd