Pith. sign in

Paper Citation Record · LEDGER

LLMind: Bio-inspired Training-free Adaptive Visual Representations for Vision-Language Models

As of 5 August 2026, this Paper Citation Record lists 58 of 58 outbound references and 2 inbound Pith citation observations for arXiv:2603.14882.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2603.14882 v2

Coverage vector

measured 58 of 58 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-15T10:31:05.743053Z

measured 60 of 60 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-08T14:44:32.242554Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-08T14:44:59.711636Z

Reference resolution

58 of 58 outbound references displayed

  • verified exact14
  • verified fuzzy44
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c32950ae-fa91-4131-b677-39c5bb600aff · outbound

This paper cites Data augmentation with noise and blur to enhance the performance of yolo7 ob- ject detection algorithm.

LLMind: Bio-inspired Training-free Adaptive Visual Representations for Vision-Language Models Data augmentation with noise and blur to enhance the performance of yolo7 ob- ject detection algorithm

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T10:39:57.694263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T10:31:05.743053Z digest=sha256:aa3179a756af1e2492ddde54ea19cef006a858ef1dc848f63b12dee198e2e827

Observation 648f6afc-2a75-4370-afd1-7968eab41f11 · outbound

This paper cites Object detection through search with a foveated visual system.PLoS com- putational biology, 13(10):e1005743.

LLMind: Bio-inspired Training-free Adaptive Visual Representations for Vision-Language Models Object detection through search with a foveated visual system.PLoS com- putational biology, 13(10):e1005743

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T10:39:57.731196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T10:31:05.743053Z digest=sha256:870a3c397c26082e69a8d038b5ad167223afa98c539c352013bd6f0eccfbbf24

Observation 98835792-3e93-48f2-889e-9ae8bbf9883e · outbound

This paper cites Flamingo: a visual language model for few-shot learning.Advances in neural information processing systems, 35:23716–23736.

LLMind: Bio-inspired Training-free Adaptive Visual Representations for Vision-Language Models Flamingo: a visual language model for few-shot learning.Advances in neural information processing systems, 35:23716–23736

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T10:39:57.734642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T10:31:05.743053Z digest=sha256:0bc9e3761e8c0cb2ee87b7b8e97c7aa0f749e15506f9cfa52016eb6cc058fd59

Observation 1dc8fe82-bf2c-473a-b38b-23e5989aab5d · outbound

This paper cites M ¨obius trans- formations revealed.Notices of the American Mathematical Society, 55(10):1226–1231.

LLMind: Bio-inspired Training-free Adaptive Visual Representations for Vision-Language Models M ¨obius trans- formations revealed.Notices of the American Mathematical Society, 55(10):1226–1231

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T10:39:57.737928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T10:31:05.743053Z digest=sha256:536963b90b5818affc0c72e36f9e85bc82f55205140acbbce98d911fc14c70bb

Observation f4ae511c-044a-404e-be6e-f055bedf5d7f · outbound

This paper cites Qwen2.5-VL Technical Report.

LLMind: Bio-inspired Training-free Adaptive Visual Representations for Vision-Language Models Qwen2.5-VL Technical Report

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-15T10:35:27.643584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T10:31:05.743053Z digest=sha256:44f43db077e353cc6a103765e5c7c720748d5f0503bbc5cae051409cad5f9062

Observation 0ffaac07-16ee-4144-bc56-e27e8abf039d · outbound

This paper cites A summary-statistic representation in peripheral vision ex- plains visual crowding.Journal of vision, 9(12):13–13.

LLMind: Bio-inspired Training-free Adaptive Visual Representations for Vision-Language Models A summary-statistic representation in peripheral vision ex- plains visual crowding.Journal of vision, 9(12):13–13

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T10:39:57.744724Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T10:31:05.743053Z digest=sha256:31f0005831fcce81fe77ded58604707b7c74f2fc353af54f61d86a7a4dd15fb2

Observation 1fb6bdf0-2b41-4cb2-a621-8e85202e8e35 · outbound

This paper cites Token Merging: Your ViT But Faster.

LLMind: Bio-inspired Training-free Adaptive Visual Representations for Vision-Language Models Token Merging: Your ViT But Faster

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-15T10:35:27.626152Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T10:31:05.743053Z digest=sha256:dd6b46ca9cd586d6cd8e4a12f4c0da14246e55d6cf7e9e7df7221c1655416ff1

Observation bf716ec9-bd95-46e7-8e6d-d5420f1b131b · outbound

This paper cites A general survey on attention mechanisms in deep learning.IEEE transactions on knowledge and data engineering, 35(4):3279–3298.

LLMind: Bio-inspired Training-free Adaptive Visual Representations for Vision-Language Models A general survey on attention mechanisms in deep learning.IEEE transactions on knowledge and data engineering, 35(4):3279–3298

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T10:39:57.727690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T10:31:05.743053Z digest=sha256:1f8c2f52b846298df4fa3fea4fd25c2549c74367d22ed020f93966ee7828f02e

Observation 27e13e15-c2fa-4f89-98ca-3ecb3745ca7b · outbound

This paper cites Convolutional neural networks develop ma- jor organizational principles of early visual cortex when en- hanced with retinal sampling.Scientific Reports, 14(1):8980.

LLMind: Bio-inspired Training-free Adaptive Visual Representations for Vision-Language Models Convolutional neural networks develop ma- jor organizational principles of early visual cortex when en- hanced with retinal sampling.Scientific Reports, 14(1):8980

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T10:39:57.741300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T10:31:05.743053Z digest=sha256:70d582626a615c43a6b65ca88f9dd6394b03b2bf51ff2a0e2ba3b3ab81a3c456

Observation f9619b07-9ced-4ece-a261-43bc2bfd7a9f · outbound

This paper cites Instructblip: Towards general-purpose vision- language models with instruction tuning.Advances in neural information processing systems, 36:49250–49267.

LLMind: Bio-inspired Training-free Adaptive Visual Representations for Vision-Language Models Instructblip: Towards general-purpose vision- language models with instruction tuning.Advances in neural information processing systems, 36:49250–49267

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T10:35:28.672840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T10:31:05.743053Z digest=sha256:b5943390d14dd55ccf75379cf57f7d96b6f39310f16e42e8d83577b21d9d38a8

Observation e38eead7-ab86-4145-8ce8-c03db7705e70 · outbound

This paper cites The representation of the visual field on the cerebral cortex in monkeys.The Journal of physiology, 159(2):203.

LLMind: Bio-inspired Training-free Adaptive Visual Representations for Vision-Language Models The representation of the visual field on the cerebral cortex in monkeys.The Journal of physiology, 159(2):203

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T10:35:28.674871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T10:31:05.743053Z digest=sha256:b394ef17438c3edd1e9afc9ff04e6afa6ae55f9ca583f70aeb6dd6edaaae7a2d

Observation 836cf222-4930-47d3-b85f-1eeb032470f2 · outbound

This paper cites PhD thesis, Indian Institute of Technology Gandhinagar.

LLMind: Bio-inspired Training-free Adaptive Visual Representations for Vision-Language Models PhD thesis, Indian Institute of Technology Gandhinagar

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T10:35:28.668644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T10:31:05.743053Z digest=sha256:0ade08e87be181aa4da34ea3441aec47c415883495003d620baf4439d5ea90d9

Observation 9bd20661-793f-4b0d-989c-b17dfc2dee5d · outbound

This paper cites Modified harris hawk optimization algorithm for mul- tilevel image thresholding.

LLMind: Bio-inspired Training-free Adaptive Visual Representations for Vision-Language Models Modified harris hawk optimization algorithm for mul- tilevel image thresholding

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T10:39:57.747620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T10:31:05.743053Z digest=sha256:4e4eb15727a2bd550358f3fe9072020d3b22c92899d2f9e594b3d8585f3cf4a3

Observation 29105450-b676-4f6f-bd49-ff593144082d · outbound

This paper cites Scribgen: generating scribble art through meta- heuristics.

LLMind: Bio-inspired Training-free Adaptive Visual Representations for Vision-Language Models Scribgen: generating scribble art through meta- heuristics

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T10:35:28.666373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T10:31:05.743053Z digest=sha256:3441cfabaf117a3cb44adf89972321213b19ce2d0ad8342bf758387895ecd55b

Observation f520765a-c60f-4bcf-9cfd-94736ae05b6e · outbound

This paper cites Emergent Properties of Foveated Perceptual Systems.

LLMind: Bio-inspired Training-free Adaptive Visual Representations for Vision-Language Models Emergent Properties of Foveated Perceptual Systems

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-15T10:35:27.617377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T10:31:05.743053Z digest=sha256:951c314fcdc1573f40c1c9ec4e34dce4a73965adc5a78405d7dc78b7fea625fa

Observation 4d725512-2d7f-40e7-a844-d7a0cf57d081 · outbound

This paper cites Image quality assessment: Unifying structure and texture similarity.IEEE transactions on pattern analysis and ma- chine intelligence, 44(5):2567–2581.

LLMind: Bio-inspired Training-free Adaptive Visual Representations for Vision-Language Models Image quality assessment: Unifying structure and texture similarity.IEEE transactions on pattern analysis and ma- chine intelligence, 44(5):2567–2581

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T10:35:28.676882Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T10:31:05.743053Z digest=sha256:46d855ab4083f5d1ccdaaaa2d60159f495d93480585dcd071f956b1c4a7d1703

Observation 5413c99f-e089-4148-b28e-19f867774949 · outbound

This paper cites Selective attention and the organization of vi- sual information.Journal of experimental psychology: Gen- eral, 113(4):501.

LLMind: Bio-inspired Training-free Adaptive Visual Representations for Vision-Language Models Selective attention and the organization of vi- sual information.Journal of experimental psychology: Gen- eral, 113(4):501

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T10:35:28.681103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T10:31:05.743053Z digest=sha256:adc210137fbbeb8080a78c74552fc07e1bf28c6a87589e7c9fbfc1d1deb9097d

Observation 5b736cc7-cc00-46f2-8146-f9de1e5efb70 · outbound

This paper cites Metamers of the ventral stream.Nature neuroscience, 14(9):1195–1201.

LLMind: Bio-inspired Training-free Adaptive Visual Representations for Vision-Language Models Metamers of the ventral stream.Nature neuroscience, 14(9):1195–1201

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T10:35:28.661486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T10:31:05.743053Z digest=sha256:ef71bfd001cba6ed79fc957a4f7780663af1c3b24f9bcc3ca5e642055ad8c0ba

Observation c080502c-a7b0-4022-8eb0-79be1878c4b9 · outbound

This paper cites The free-energy principle: a unified brain the- ory?Nature reviews neuroscience, 11(2):127–138.

LLMind: Bio-inspired Training-free Adaptive Visual Representations for Vision-Language Models The free-energy principle: a unified brain the- ory?Nature reviews neuroscience, 11(2):127–138

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T10:35:28.659117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T10:31:05.743053Z digest=sha256:56d75c88b32719feb6c32efd2076face684c99645d1546e303c7b265f93d6199

Observation ed7dd5ed-17e5-4800-a57c-90329ed1a442 · outbound

This paper cites FeatUp: A Model-Agnostic Framework for Features at Any Resolution.

LLMind: Bio-inspired Training-free Adaptive Visual Representations for Vision-Language Models FeatUp: A Model-Agnostic Framework for Features at Any Resolution

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-15T10:35:27.608930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T10:31:05.743053Z digest=sha256:ab2365110d972969745dcba8cca277e4c6c3ca13d3f4a4f1a8002f713b21288d

Observation 749041e1-ea54-4138-ba47-b955f65917e9 · outbound

This paper cites Mini-internvl: a flexible-transfer pocket multi-modal model with 5% parameters and 90% perfor- mance.Visual Intelligence, 2(1):32.

LLMind: Bio-inspired Training-free Adaptive Visual Representations for Vision-Language Models Mini-internvl: a flexible-transfer pocket multi-modal model with 5% parameters and 90% perfor- mance.Visual Intelligence, 2(1):32

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T10:39:57.762710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T10:31:05.743053Z digest=sha256:a12af6679e58bbf714577c06c0c5cf3bccb6baba87818a878e3be8b78cca3030

Observation 52dbf0eb-e94f-4dd0-bd72-c73bb32557bf · outbound

This paper cites Vari- able resolution improves visual question answering under a limited pixel budget.

LLMind: Bio-inspired Training-free Adaptive Visual Representations for Vision-Language Models Vari- able resolution improves visual question answering under a limited pixel budget

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T10:35:28.650273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T10:31:05.743053Z digest=sha256:5b792b52644232033da5ba353369442fc8c92ba10b01f3a0dcfa93f01d502d04

Observation 990b23ff-9a13-40d3-bd34-b84cd34ab078 · outbound

This paper cites See- ing more with less: Human-like representations in vision models.

LLMind: Bio-inspired Training-free Adaptive Visual Representations for Vision-Language Models See- ing more with less: Human-like representations in vision models

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T10:35:28.654698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T10:31:05.743053Z digest=sha256:c7d762fe8ec376df3bc6141b88df71d90d8226560cb7eabc970622517e5617dd

Observation 07a056aa-f506-4a0a-9fdc-862517fda58c · outbound

This paper cites Making the v in vqa matter: Elevating the role of image understanding in visual question answer- ing.

LLMind: Bio-inspired Training-free Adaptive Visual Representations for Vision-Language Models Making the v in vqa matter: Elevating the role of image understanding in visual question answer- ing

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T10:39:57.713950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T10:31:05.743053Z digest=sha256:4344652d3e795575c8b0a5a4721c775be32e2d1f5f37f3d56a4de88b66b50ee1

Observation 0466e7cf-1a37-4ca0-a7a6-6f679a77bec9 · outbound

This paper cites Lvis: A dataset for large vocabulary instance segmentation.

LLMind: Bio-inspired Training-free Adaptive Visual Representations for Vision-Language Models Lvis: A dataset for large vocabulary instance segmentation

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T10:35:28.646080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T10:31:05.743053Z digest=sha256:1e0eb0289b4c71ced03137cb6a428bf51319c11e96f7d377a13a5137a958c84a

Observation a5c62fa2-5f0a-4ae3-b774-f519f8053f40 · outbound

This paper cites A survey on vision transformer.

LLMind: Bio-inspired Training-free Adaptive Visual Representations for Vision-Language Models A survey on vision transformer

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T10:35:28.643977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T10:31:05.743053Z digest=sha256:c6de21733a11eb5bd8d4366b593dc3757694f3040c2372582630648231caa30c

Observation 82eb9942-4941-40b6-9882-1069f6e1fec9 · outbound

This paper cites Coco-periph: bridging the gap between human and machine perception in the periphery.

LLMind: Bio-inspired Training-free Adaptive Visual Representations for Vision-Language Models Coco-periph: bridging the gap between human and machine perception in the periphery

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T10:35:28.648058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T10:31:05.743053Z digest=sha256:7217db629c9092d336c7abfb6e5303d19994992d2bc5afdf54f359475f92d84b

Observation 14c681cd-a103-4d04-a332-dac975a1aea9 · outbound

This paper cites Eye movements in natural behavior.Trends in cognitive sciences, 9(4):188–194.

LLMind: Bio-inspired Training-free Adaptive Visual Representations for Vision-Language Models Eye movements in natural behavior.Trends in cognitive sciences, 9(4):188–194

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T10:35:28.652701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T10:31:05.743053Z digest=sha256:9fd4a2c31760027580956b7409d1aae15990915d2c39301f5d20679d1a04759d

Observation 79981c56-12b0-4e38-9522-7e19687406cb · outbound

This paper cites Vi- sual prompt tuning.

LLMind: Bio-inspired Training-free Adaptive Visual Representations for Vision-Language Models Vi- sual prompt tuning

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T10:35:28.642204Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T10:31:05.743053Z digest=sha256:f821fd37437f82d8da33c67789f7123c2fbdf52817cb8ae7bf75cb940b581fe0

Observation 929a1275-463c-42d0-9969-634425b509b0 · outbound

This paper cites Transformers in vision: A survey.ACM computing surveys (CSUR), 54(10s):1–41.

LLMind: Bio-inspired Training-free Adaptive Visual Representations for Vision-Language Models Transformers in vision: A survey.ACM computing surveys (CSUR), 54(10s):1–41

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T10:35:28.640412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T10:31:05.743053Z digest=sha256:4842d88ba5e6b3ccd89ca7ed851e306b7077efda59af81f684c24a4c6f4fd838

Observation d99382e1-c4b5-440e-8e40-64155488d60f · outbound

This paper cites Foveation in the Era of Deep Learning.

LLMind: Bio-inspired Training-free Adaptive Visual Representations for Vision-Language Models Foveation in the Era of Deep Learning

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-15T10:35:27.662926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T10:31:05.743053Z digest=sha256:f9371e5c36c178f4c8bb7d05e4a7a4d9e35aa636121f9456aa709491c48241bc

Observation d122a807-d21f-4a6a-bf84-9729c95af1f8 · outbound

This paper cites Oxford Uni- versity Press.

LLMind: Bio-inspired Training-free Adaptive Visual Representations for Vision-Language Models Oxford Uni- versity Press

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T10:35:28.670665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T10:31:05.743053Z digest=sha256:bce410a7c05e60431a3076a387dfc31711e2156b671059790c82ad1798d4161c

Observation 7ca8e6f7-f59b-4d3c-915e-01d66a709c4b · outbound

This paper cites SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension.

LLMind: Bio-inspired Training-free Adaptive Visual Representations for Vision-Language Models SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-05-15T10:35:27.613119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T10:31:05.743053Z digest=sha256:ae46ec8ac89429eb7056087126e53080065b673f669ef107f19bdb3374b9198e

Observation 99321951-9ab3-4bcd-938e-2702b5060799 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

LLMind: Bio-inspired Training-free Adaptive Visual Representations for Vision-Language Models LLaVA-OneVision: Easy Visual Task Transfer

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-05-15T10:35:27.647915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T10:31:05.743053Z digest=sha256:28f2e093b9dace098b021b8d33824b9d25cfeb8b2e77122166c1f7b23875d533

Observation 4cc72545-f925-4f5b-9606-645e278c9179 · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

LLMind: Bio-inspired Training-free Adaptive Visual Representations for Vision-Language Models Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T10:35:28.638366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T10:31:05.743053Z digest=sha256:43d2ee2d891a7768addfed6a70356db11aa9ade236852a884008e7be7662eee5

Observation 5063c342-afa5-446b-9dd7-63d8c0d2ed08 · outbound

This paper cites DeepSeek-VL: Towards Real-World Vision-Language Understanding.

LLMind: Bio-inspired Training-free Adaptive Visual Representations for Vision-Language Models DeepSeek-VL: Towards Real-World Vision-Language Understanding

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-05-15T10:35:27.621585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T10:31:05.743053Z digest=sha256:cc91a379ee87564bec05342e5e48b0fa1debadcc8d60bd7ee14899aabf06456a

Observation 30e190e7-0b4c-42d4-871d-2a753af1ae27 · outbound

This paper cites Bio- logically inspired deep learning model for efficient foveal- peripheral vision.Frontiers in Computational Neuroscience, 15:746204.

LLMind: Bio-inspired Training-free Adaptive Visual Representations for Vision-Language Models Bio- logically inspired deep learning model for efficient foveal- peripheral vision.Frontiers in Computational Neuroscience, 15:746204

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T10:35:28.656836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T10:31:05.743053Z digest=sha256:07b4afef9d129f3a9843cbea4c6e9c962b3ff62a570cd9c4deec107974aa5427

Observation 7b332d7e-6c2c-4cf3-ab75-c75b15888d8c · outbound

This paper cites Mind Meets Space: Rethinking Agentic Spatial Intelligence from a Neuroscience-inspired Perspective.

LLMind: Bio-inspired Training-free Adaptive Visual Representations for Vision-Language Models Mind Meets Space: Rethinking Agentic Spatial Intelligence from a Neuroscience-inspired Perspective

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-15T10:35:27.635083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T10:31:05.743053Z digest=sha256:d60e72154b9b928eb7936902ffa0dbafd551702195b137c317e36a21668fcb72

Observation a03270dc-468f-497d-95dc-4e01ce2af4d9 · outbound

This paper cites SmolVLM: Redefining small and efficient multimodal models.

LLMind: Bio-inspired Training-free Adaptive Visual Representations for Vision-Language Models SmolVLM: Redefining small and efficient multimodal models

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-05-15T10:35:27.604164Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T10:31:05.743053Z digest=sha256:b1596604913909c2c101d1111cb34a78bae45cef49437cd0f4470a1ee39d90a9

Observation 3e1398e7-de42-4952-a76d-a94d53d1d51e · outbound

This paper cites Peripheral vision transformer.Advances in Neural Informa- tion Processing Systems, 35:32097–32111.

LLMind: Bio-inspired Training-free Adaptive Visual Representations for Vision-Language Models Peripheral vision transformer.Advances in Neural Informa- tion Processing Systems, 35:32097–32111

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T10:35:28.664296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T10:31:05.743053Z digest=sha256:aad8e856e701e34f51321daebfa8616b98e6959d19de74a8d03789794d4a4d10

Observation ff5e0932-77bb-4a0e-8bf0-8a36035ca3a8 · outbound

This paper cites The geometry of m ¨obius transformations.

LLMind: Bio-inspired Training-free Adaptive Visual Representations for Vision-Language Models The geometry of m ¨obius transformations

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T10:39:57.756118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T10:31:05.743053Z digest=sha256:b20155edab29c046381cff75c1cd22fc5eeb39aa30925a5217453d9bdbf82d82

Observation 5b08aeba-4620-41e4-a457-0e53e67a8a2f · outbound

This paper cites Learning to search for and detect objects in foveal images using deep learning.

LLMind: Bio-inspired Training-free Adaptive Visual Representations for Vision-Language Models Learning to search for and detect objects in foveal images using deep learning

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T10:39:57.753523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T10:31:05.743053Z digest=sha256:85379206a6bb303104eeef283167aa6a087b2d1924ae51126ba2f9738558f85d

Observation 52f67fdc-44b3-4d3c-bb24-706a049ab5cb · outbound

This paper cites Human peripheral blur is optimal for object recognition.Vision research, 200: 108083.

LLMind: Bio-inspired Training-free Adaptive Visual Representations for Vision-Language Models Human peripheral blur is optimal for object recognition.Vision research, 200: 108083

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T10:39:57.710594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T10:31:05.743053Z digest=sha256:c69acbaec424aeb783c8c59828860ebbe8f1c695aeb03b2e923d2afe0abd301a

Observation a94cd915-bce3-4f97-b45c-40659c023607 · outbound

This paper cites Dynamicvit: Efficient vision transformers with dynamic token sparsification.Advances in neural information processing systems, 34:13937–13949.

LLMind: Bio-inspired Training-free Adaptive Visual Representations for Vision-Language Models Dynamicvit: Efficient vision transformers with dynamic token sparsification.Advances in neural information processing systems, 34:13937–13949

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T10:39:57.750484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T10:31:05.743053Z digest=sha256:253e0982e2425a07899ada3d4a4f45c6a9410919773c772bc328eab822df2445

Observation c74410da-f002-4040-a94e-d837bdc379a8 · outbound

This paper cites Biologically inspired image sampling for electronic eye.

LLMind: Bio-inspired Training-free Adaptive Visual Representations for Vision-Language Models Biologically inspired image sampling for electronic eye

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T10:39:57.723571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T10:31:05.743053Z digest=sha256:726331bcc356e549a35246da4019ab5b0e6ee65cb0140b14565205a14319ef2c

Observation eaa78c98-22e2-47ad-9a4c-19cd4889d51e · outbound

This paper cites Robotic materials with bioinspired microstructures for high sensitivity and fast ac- tuation.Advanced Science, 13(15):e09739.

LLMind: Bio-inspired Training-free Adaptive Visual Representations for Vision-Language Models Robotic materials with bioinspired microstructures for high sensitivity and fast ac- tuation.Advanced Science, 13(15):e09739

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T10:39:57.717368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T10:31:05.743053Z digest=sha256:d7805070a3ca08fc268abe3b67e8870931100fd34609262e9a812e166b97bbae

Observation 0218b050-ffb4-4e88-b015-d151fb2e6cf9 · outbound

This paper cites A-okvqa: A benchmark for visual question answering using world knowl- edge.

LLMind: Bio-inspired Training-free Adaptive Visual Representations for Vision-Language Models A-okvqa: A benchmark for visual question answering using world knowl- edge

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T10:39:57.706937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T10:31:05.743053Z digest=sha256:177845ecf231ac4ab1ec35a2d2c6b4902d5a270c9e092a1ea393812a06da6028

Observation d7534fa2-d532-452e-90a6-86752c1da381 · outbound

This paper cites Behind the Machine's Gaze: Neural Networks with Biologically-inspired Constraints Exhibit Human-like Visual Attention.

LLMind: Bio-inspired Training-free Adaptive Visual Representations for Vision-Language Models Behind the Machine's Gaze: Neural Networks with Biologically-inspired Constraints Exhibit Human-like Visual Attention

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-15T10:35:27.658311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T10:31:05.743053Z digest=sha256:da94c56d1de74d03a0d05cefb79091603edf7b90ad2b6f4910b297b07a2fdacb

Observation f3921aa6-5a35-43f5-ab29-7575a8922137 · outbound

This paper cites Multivariate stochastic approximation using a simultaneous perturbation gradient approximation.IEEE transactions on automatic control, 37(3):332–341.

LLMind: Bio-inspired Training-free Adaptive Visual Representations for Vision-Language Models Multivariate stochastic approximation using a simultaneous perturbation gradient approximation.IEEE transactions on automatic control, 37(3):332–341

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T10:35:28.678951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T10:31:05.743053Z digest=sha256:14dd29987b3dc6b9fe522bb69c98892607a060dc3966d6fab786c279527cce09

Observation af1f6430-0ad6-415a-82e8-961a673bfa71 · outbound

This paper cites Pe- ripheral vision and pattern recognition: A review.Journal of vision, 11(5):13–13.

LLMind: Bio-inspired Training-free Adaptive Visual Representations for Vision-Language Models Pe- ripheral vision and pattern recognition: A review.Journal of vision, 11(5):13–13

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T10:39:57.759239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T10:31:05.743053Z digest=sha256:65c53ed98736b3fb65f0681468db80b247f80305ae299aa8804d8c874f1f4ab5

Observation 41923778-ee58-4d55-a2a5-73b6a9255a3f · outbound

This paper cites Central and peripheral vision for scene recognition: A neurocomputational model- ing exploration.Journal of vision, 17(4):9–9.

LLMind: Bio-inspired Training-free Adaptive Visual Representations for Vision-Language Models Central and peripheral vision for scene recognition: A neurocomputational model- ing exploration.Journal of vision, 17(4):9–9

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T10:39:57.697699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T10:31:05.743053Z digest=sha256:7aaf6fcb6938bfc1410f72c2865f235024d303b2cd1b857931fd58a00a90a174

Observation d9ea73e5-70da-456b-bb51-9e6a76e7037c · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

LLMind: Bio-inspired Training-free Adaptive Visual Representations for Vision-Language Models Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 52

Resolution
verified exact
local_arxiv, observed 2026-05-15T10:35:27.630268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T10:31:05.743053Z digest=sha256:074c72ea4e415f1ef596a950ef687c2b4e1f38600b4a5badeb4b3ba9bfd4087d

Observation 88e62a00-72a6-4ce4-89da-d692cf9e9866 · outbound

This paper cites Minilm: Deep self-attention distillation for task-agnostic compression of pre-trained transformers.Ad- vances in neural information processing systems, 33:5776– 5788.

LLMind: Bio-inspired Training-free Adaptive Visual Representations for Vision-Language Models Minilm: Deep self-attention distillation for task-agnostic compression of pre-trained transformers.Ad- vances in neural information processing systems, 33:5776– 5788

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T10:39:57.720624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T10:31:05.743053Z digest=sha256:dbb496bc566acec0ec4ec8ebfe408896c29a148853b6f603020c751540c586ba

Observation 573345b3-4ac8-47be-95dd-b4a9928691c7 · outbound

This paper cites Emulating human-like adaptive vision for efficient and flexible machine visual perception.Nature Machine In- telligence, pages 1–19.

LLMind: Bio-inspired Training-free Adaptive Visual Representations for Vision-Language Models Emulating human-like adaptive vision for efficient and flexible machine visual perception.Nature Machine In- telligence, pages 1–19

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T10:39:57.703051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T10:31:05.743053Z digest=sha256:6aad66c9568ba205d06030e3e7c0daa0aea8829bc5185bed4b28f9f8ddaba891

Observation 01a0c6a9-d91f-4ce8-960d-4fdeeda6d0d8 · outbound

This paper cites Controlmllm: Training-free visual prompt learning for multimodal large language models.Advances in Neural Information Processing Systems, 37:45206–45234.

LLMind: Bio-inspired Training-free Adaptive Visual Representations for Vision-Language Models Controlmllm: Training-free visual prompt learning for multimodal large language models.Advances in Neural Information Processing Systems, 37:45206–45234

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T10:39:57.702339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T10:31:05.743053Z digest=sha256:a40163bdc0a40b7e3fa976fd98cde9ea8e4a53b401d485fd9d8a64c7ffb8301e

Observation 7cba3786-890d-45cd-ae9c-bc4f6b09d91e · outbound

This paper cites Qwen3 Technical Report.

LLMind: Bio-inspired Training-free Adaptive Visual Representations for Vision-Language Models Qwen3 Technical Report

Reference 56

Resolution
verified exact
local_arxiv, observed 2026-05-15T10:35:27.639521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T10:31:05.743053Z digest=sha256:0175e5038946b150cd00fc4f3ae59683b4a60feb1b55b70e67a7c2d45218ebde

Observation ba856bf3-d05e-4a5b-8f09-33256ac6d2d3 · outbound

This paper cites Ferret: Refer and Ground Anything Anywhere at Any Granularity.

LLMind: Bio-inspired Training-free Adaptive Visual Representations for Vision-Language Models Ferret: Refer and Ground Anything Anywhere at Any Granularity

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-05-15T11:29:05.422888Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T10:31:05.743053Z digest=sha256:6dbc40913f4a7750e4ff1d0914bd1b443e41f00f018a2b89e2c47f4e3207f6ac

Observation 7e47f42b-2de2-4bdb-9986-2f98378a40c9 · outbound

This paper cites Vsi: A visual saliency-induced index for perceptual image quality assess- ment.IEEE Transactions on Image processing, 23(10): 4270–4281.

LLMind: Bio-inspired Training-free Adaptive Visual Representations for Vision-Language Models Vsi: A visual saliency-induced index for perceptual image quality assess- ment.IEEE Transactions on Image processing, 23(10): 4270–4281

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T10:39:57.685691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T10:31:05.743053Z digest=sha256:5cb9510404f5da67ab15a5440bc7bade3debf1a6461a81e9df8b2acb9c4dec5c

Pith citing papers

Observation bc148bdd-efd0-49e6-9c82-564e8ed3f352 · inbound

Toward Native Multimodal Modeling: A Roadmap cites this paper.

Toward Native Multimodal Modeling: A Roadmap LLMind: Bio-inspired Training-free Adaptive Visual Representations for Vision-Language Models

Reference 261

Resolution
verified exact
local_arxiv, observed 2026-06-29T23:04:01.864944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-29T22:58:38.610609Z digest=sha256:b5354babe553958741ed680d751c40c4bbed69c24f71738acee426f4d1ab8b6c

Observation 2ca7f7bd-a232-4a5d-b24c-c506c11375b3 · inbound

EAGOR: Embodied Reasoning in Omni-direction cites this paper.

EAGOR: Embodied Reasoning in Omni-direction LLMind: Bio-inspired Training-free Adaptive Visual Representations for Vision-Language Models

Reference 53

Resolution
verified exact
local_arxiv, observed 2026-07-08T14:44:59.712899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-08T14:44:32.242554Z digest=sha256:063694335a058e0e3c533372ed488a821258c465060ee7a272393a057aaa928f