Pith. sign in

Paper Citation Record · LEDGER

Beyond Sight: Towards Cognitive Alignment in LVLM via Enriched Visual Knowledge

As of 15 August 2026, this Paper Citation Record lists 57 of 57 outbound references and 1 inbound Pith citation observation for arXiv:2411.16824.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.16824 v1

Coverage vector

measured 57 of 57 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T12:59:46.961138Z

measured 58 of 58 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T10:23:12.580211Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T10:23:12.613702Z

Reference resolution

57 of 57 outbound references displayed

  • verified exact0
  • verified fuzzy21
  • unresolved34
  • parse uncertain1
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a0695d83-0a85-4631-9a85-397e1c593e00 · outbound

This paper cites Llama 3 model card.

Beyond Sight: Towards Cognitive Alignment in LVLM via Enriched Visual Knowledge Llama 3 model card

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T12:59:46.649890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:59:46.649890Z digest=sha256:20737409d750ef0503aae04625f6b06e9f3febe973bbe90df1691d552baf0e45

Observation 08c112a0-3e16-4516-99e1-15bd5f9f681b · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

Beyond Sight: Towards Cognitive Alignment in LVLM via Enriched Visual Knowledge Flamingo: a visual language model for few-shot learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T12:59:46.656025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:59:46.656025Z digest=sha256:2874eb2ed7184644b9331ba81a9b34d00372b2c71aef5e163204c1d8ecf17011

Observation bfcb16bb-701d-4870-a262-d996e2aba22a · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Beyond Sight: Towards Cognitive Alignment in LVLM via Enriched Visual Knowledge Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T12:59:46.661613Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:59:46.661613Z digest=sha256:bd3d6ed0d9d23efa175e0312f081841dbe85f41ab9e77d39c43a83b6a135f6da

Observation eaf489e2-279c-4785-a40d-7a05b0ce8825 · outbound

This paper cites Improving fine-grained understanding in image-text pre-training.

Beyond Sight: Towards Cognitive Alignment in LVLM via Enriched Visual Knowledge Improving fine-grained understanding in image-text pre-training

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T12:59:46.667936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:59:46.667936Z digest=sha256:ec4c7923700148311734bae2bc2a8ddbd7879397168ff4f3ff9900286bf29488

Observation 4a056aa9-3c99-4753-89d4-59e941578d09 · outbound

This paper cites Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic.

Beyond Sight: Towards Cognitive Alignment in LVLM via Enriched Visual Knowledge Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T12:59:46.673456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:59:46.673456Z digest=sha256:1f0eb0cd58bfda79d473f69c8b4b464651115f354d582e4184a3a86e588b2cc1

Observation f343a77a-991f-4c2d-ae99-e750c5b2b9b7 · outbound

This paper cites InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning.

Beyond Sight: Towards Cognitive Alignment in LVLM via Enriched Visual Knowledge InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T12:59:46.679110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:59:46.679110Z digest=sha256:96d18925ed7c518a4da5e282415bbe9d2d95e01b2833bcc4bef03e3d8816ba59

Observation fb9168b0-f1d2-4eda-83df-06b59c5464d1 · outbound

This paper cites InternLM-XComposer2-4KHD: A Pioneering Large Vision-Language Model Handling Resolutions from 336 Pixels to 4K HD.

Beyond Sight: Towards Cognitive Alignment in LVLM via Enriched Visual Knowledge InternLM-XComposer2-4KHD: A Pioneering Large Vision-Language Model Handling Resolutions from 336 Pixels to 4K HD

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T12:59:46.684673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:59:46.684673Z digest=sha256:5accb44c32580d41bd1cd24a33855ad21b1b94793f0ac46f7b304b5740c3fb94

Observation e8910bbd-b644-42c1-985d-6c6d9f941408 · outbound

This paper cites Does Fine-Tuning LLMs on New Knowledge Encourage Hallucinations?.

Beyond Sight: Towards Cognitive Alignment in LVLM via Enriched Visual Knowledge Does Fine-Tuning LLMs on New Knowledge Encourage Hallucinations?

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T12:59:46.691039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:59:46.691039Z digest=sha256:185d24330e28522299a8bfe42640e52a5b1f4ea68452ba689bc552304a6667df

Observation 7bafba62-0bf5-4efd-a4ca-2a7e7a3dae25 · outbound

This paper cites Bard, 2023.

Beyond Sight: Towards Cognitive Alignment in LVLM via Enriched Visual Knowledge Bard, 2023

Reference 9

Resolution
parse uncertain
no resolver link, observed 2026-08-12T12:59:46.696337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:59:46.696337Z digest=sha256:137baabc032ddac355de1d5c4ec09e6374cffbb24bcffe9baaf2ded501b7cac6

Observation 6f67d8a9-a55a-40d9-a70b-aeabf8ff9c47 · outbound

This paper cites Making the V in VQA matter: El- evating the role of image understanding in visual question answering.

Beyond Sight: Towards Cognitive Alignment in LVLM via Enriched Visual Knowledge Making the V in VQA matter: El- evating the role of image understanding in visual question answering

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:59:48.065563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T12:59:46.702313Z digest=sha256:071f5f2676a087f57af4cd4462538f557d2dd4fd3e087f1661982ca74a1a2ab6

Observation 5b567098-450f-4c77-a295-d70ed847689a · outbound

This paper cites GQA: A new dataset for real-world visual reasoning and compositional question answering.

Beyond Sight: Towards Cognitive Alignment in LVLM via Enriched Visual Knowledge GQA: A new dataset for real-world visual reasoning and compositional question answering

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:59:48.047149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T12:59:46.708280Z digest=sha256:adad6b9d48d62ee009ec7e3c82761de6513628253cf4d482f53e305368c7047f

Observation a6228346-387d-4473-a46f-d252efea1e45 · outbound

This paper cites BRAVE: Broadening the visual encoding of vision-language models.

Beyond Sight: Towards Cognitive Alignment in LVLM via Enriched Visual Knowledge BRAVE: Broadening the visual encoding of vision-language models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T12:59:46.713507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:59:46.713507Z digest=sha256:aeebc3f1d484386759d15b8e4241cd7236e43987c85b10865637574445e491bf

Observation 7c77ee76-4c8b-407a-9d52-51a893626f05 · outbound

This paper cites Referitgame: Referring to objects in pho- tographs of natural scenes.

Beyond Sight: Towards Cognitive Alignment in LVLM via Enriched Visual Knowledge Referitgame: Referring to objects in pho- tographs of natural scenes

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:59:48.028406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T12:59:46.719540Z digest=sha256:68cc8a050f7d8ec3b17d6abbff8761050e6ecb83423d71590ad705d9d7322412

Observation 40b7d9a5-72b7-41a7-8620-2be600bc6ff9 · outbound

This paper cites Visual genome: Connecting language and vision using crowdsourced dense image annotations.

Beyond Sight: Towards Cognitive Alignment in LVLM via Enriched Visual Knowledge Visual genome: Connecting language and vision using crowdsourced dense image annotations

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:59:48.010433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T12:59:46.724572Z digest=sha256:569f9e38b5bbbbed0764fc9890e02d3a8e21335cbe7f322c3252d59960574fa2

Observation af80f546-e2ee-458d-95ac-5b6f4a554bf2 · outbound

This paper cites What matters when building vision-language models?.

Beyond Sight: Towards Cognitive Alignment in LVLM via Enriched Visual Knowledge What matters when building vision-language models?

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T12:59:46.729574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:59:46.729574Z digest=sha256:7f3109c8563f8014cb7dcb2e362fdb556749134cc410ee68ba278264057b565a

Observation 206f8b15-97c3-4580-b4fa-10851a71432e · outbound

This paper cites BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models.

Beyond Sight: Towards Cognitive Alignment in LVLM via Enriched Visual Knowledge BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T12:59:46.734942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:59:46.734942Z digest=sha256:15136b507eb5db1b6e82d4b82b2164690e71d27509db6edfb73174a457a771f2

Observation a31b0483-7240-4424-97f5-13afc61b19f1 · outbound

This paper cites Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models.

Beyond Sight: Towards Cognitive Alignment in LVLM via Enriched Visual Knowledge Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T12:59:46.740462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:59:46.740462Z digest=sha256:5fcb0dd5e96c44e2730232a22e54b4a33da85095ece5c86d3b433828cfcc6247

Observation d7cabda7-85ed-4be5-9778-0f642e98ae99 · outbound

This paper cites SPHINX-X: Scaling Data and Parameters for a Family of Multi-modal Large Language Models.

Beyond Sight: Towards Cognitive Alignment in LVLM via Enriched Visual Knowledge SPHINX-X: Scaling Data and Parameters for a Family of Multi-modal Large Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T12:59:46.747182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:59:46.747182Z digest=sha256:08798183d34d3dc4e6ae9ec51a14fadc1b1ff64da117ae6e0257450c2d241af7

Observation f20cbc10-451d-4e7f-b18f-41c17b56841b · outbound

This paper cites Improved Baselines with Visual Instruction Tuning.

Beyond Sight: Towards Cognitive Alignment in LVLM via Enriched Visual Knowledge Improved Baselines with Visual Instruction Tuning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T12:59:46.752949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:59:46.752949Z digest=sha256:82c5b4ad88903e8e9329f611f1436ea52f1c437c84442c7433900614be4d7b8f

Observation 2ceadc13-e658-4322-8c20-7c5dd733449c · outbound

This paper cites Visual instruction tuning.

Beyond Sight: Towards Cognitive Alignment in LVLM via Enriched Visual Knowledge Visual instruction tuning

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:59:47.992321Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T12:59:46.758833Z digest=sha256:bf33374b8431b4f937cf6f31f3dfdeca6dcad311c15471d1bccc680f0b072eb3

Observation 90a69544-958d-486e-a519-5d9d6b57ce32 · outbound

This paper cites Decoupled weight decay regularization.

Beyond Sight: Towards Cognitive Alignment in LVLM via Enriched Visual Knowledge Decoupled weight decay regularization

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T12:59:46.763868Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:59:46.763868Z digest=sha256:823f5721bb6524aa9e9b2f93bf521d6d0987e7b093e41f51611f0553dda066cd

Observation 785dbdae-3e40-491d-aced-212468d13fc8 · outbound

This paper cites DeepSeek-VL: Towards Real-World Vision-Language Understanding.

Beyond Sight: Towards Cognitive Alignment in LVLM via Enriched Visual Knowledge DeepSeek-VL: Towards Real-World Vision-Language Understanding

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T12:59:46.769720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:59:46.769720Z digest=sha256:98ab94080443dae9da073e0140324737202d3dc8b3043d4397e16c7a86bea8fb

Observation b930b0c0-a931-402a-a0e0-ecbb4ee55ca2 · outbound

This paper cites Generation and comprehension of unambiguous object descriptions.

Beyond Sight: Towards Cognitive Alignment in LVLM via Enriched Visual Knowledge Generation and comprehension of unambiguous object descriptions

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T12:59:46.776098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:59:46.776098Z digest=sha256:f81a53a504447897b31d69a8428ca2867693876cdd9224854e410ef19b0712bb

Observation 1df21175-905e-4737-94b2-17a3c7600c80 · outbound

This paper cites OK-VQA: A visual question answering benchmark requiring external knowledge.

Beyond Sight: Towards Cognitive Alignment in LVLM via Enriched Visual Knowledge OK-VQA: A visual question answering benchmark requiring external knowledge

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:59:47.950864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T12:59:46.781351Z digest=sha256:dbae7c9ccdbf1d20dcfdf21d80b91b1b6f90c0acb4f91f9e8e5259319763b31a

Observation 3f028d67-e504-41d0-889c-0bbf12920656 · outbound

This paper cites OCR-VQA: Visual question answer- ing by reading text in images.

Beyond Sight: Towards Cognitive Alignment in LVLM via Enriched Visual Knowledge OCR-VQA: Visual question answer- ing by reading text in images

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:59:47.931349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T12:59:46.786462Z digest=sha256:b6091df5603c8e0ea6d96e306d946b38282b43fc89ca4bcafd15ae6be824cb06

Observation 27186e34-2e09-444b-bf47-f7f97774a9df · outbound

This paper cites Gpt-4 technical report, 2023.

Beyond Sight: Towards Cognitive Alignment in LVLM via Enriched Visual Knowledge Gpt-4 technical report, 2023

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:59:47.914084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T12:59:46.792118Z digest=sha256:4fe2f100a05f4f94f83a49f2789b61fe95ea20bb6f9f2fee251d8efde2202c8a

Observation dd7db9c0-3ba7-42f8-922a-a90bf145094b · outbound

This paper cites GPT-4o System Card, 2024.

Beyond Sight: Towards Cognitive Alignment in LVLM via Enriched Visual Knowledge GPT-4o System Card, 2024

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:59:47.896246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T12:59:46.797391Z digest=sha256:2ffc32e566f313bf60fdaaab5f3e8ea5070b7a4073bcdd5183046ea546a944d3

Observation 664e6cba-6119-4a20-8a07-33b712634612 · outbound

This paper cites Instruction Tuning with GPT-4.

Beyond Sight: Towards Cognitive Alignment in LVLM via Enriched Visual Knowledge Instruction Tuning with GPT-4

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T12:59:46.802681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:59:46.802681Z digest=sha256:0ac1117fa75d5ff4e205470d4692d5bce61529843a54c825890ddb0c69c2c8a5

Observation 2f5af47a-8438-41c3-a1f1-5726e5b82249 · outbound

This paper cites Learning transferable visual models from natural language supervision, 2021.

Beyond Sight: Towards Cognitive Alignment in LVLM via Enriched Visual Knowledge Learning transferable visual models from natural language supervision, 2021

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:59:47.877764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T12:59:46.807971Z digest=sha256:331032678a181a371f385da045d8ad0f8267b5eccda0997497db604ea62e2111

Observation f2379740-2995-44d1-8a84-7211cdd51519 · outbound

This paper cites A-OKVQA: A benchmark for visual question answering using world knowl- edge.

Beyond Sight: Towards Cognitive Alignment in LVLM via Enriched Visual Knowledge A-OKVQA: A benchmark for visual question answering using world knowl- edge

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:59:47.860636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T12:59:46.813150Z digest=sha256:c8ce3fe1291443da0a15408c915f62504ed69499197f0fba76d7a430d29ac074

Observation a97eb258-fb28-49d7-893f-53b94bb17d46 · outbound

This paper cites Conceptual captions: A cleaned, hypernymed, im- age alt-text dataset for automatic image captioning.

Beyond Sight: Towards Cognitive Alignment in LVLM via Enriched Visual Knowledge Conceptual captions: A cleaned, hypernymed, im- age alt-text dataset for automatic image captioning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T12:59:46.818774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:59:46.818774Z digest=sha256:0d42bc307196ad8fafc7cc3691e6fce957c9cf313d4b9367ea3bd2a5b8f6c8ef

Observation 57d375b8-60f2-45ec-b788-291bb1f1c070 · outbound

This paper cites When do we not need larger vision models? In European Conference on Computer Vision, pages 444–462.

Beyond Sight: Towards Cognitive Alignment in LVLM via Enriched Visual Knowledge When do we not need larger vision models? In European Conference on Computer Vision, pages 444–462

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:59:47.831721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T12:59:46.824079Z digest=sha256:56cc92f7aeaa138764b435eadfdb0c509443874324d022ea3c25b4f6e1fdda9d

Observation f76eff6c-0d61-4de5-873a-f521408de1a1 · outbound

This paper cites Textcaps: a dataset for image captioning with reading comprehension.

Beyond Sight: Towards Cognitive Alignment in LVLM via Enriched Visual Knowledge Textcaps: a dataset for image captioning with reading comprehension

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:59:47.813480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T12:59:46.829293Z digest=sha256:aa177083a25202af4605fce84fbc447a91d4f77e3a84294a39135a0dcd498236

Observation 98f624ac-cd17-420f-841f-6603d7b795cb · outbound

This paper cites EVA-CLIP: Improved Training Techniques for CLIP at Scale.

Beyond Sight: Towards Cognitive Alignment in LVLM via Enriched Visual Knowledge EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T12:59:46.834831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:59:46.834831Z digest=sha256:afa470c2e6d41a6ef418d3c3e6e40d32d0b102ffc34b475c1b39f95670cb6e1d

Observation b915ea51-373b-4822-965c-37ef940d6ae0 · outbound

This paper cites When are lemons purple? the concept association bias of vision-language models.

Beyond Sight: Towards Cognitive Alignment in LVLM via Enriched Visual Knowledge When are lemons purple? the concept association bias of vision-language models

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:59:47.793733Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T12:59:46.840449Z digest=sha256:78dc6bec08a3118f349d9ec0e55748d3cc2946919b64b99dfc84fb22633ea485

Observation 92ac9bfa-b275-4876-a148-4bcdad1dad42 · outbound

This paper cites Qwen2.5: A party of foundation models, 2024.

Beyond Sight: Towards Cognitive Alignment in LVLM via Enriched Visual Knowledge Qwen2.5: A party of foundation models, 2024

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:59:47.775962Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T12:59:46.846151Z digest=sha256:a847c0d3a1d09ce17c341ecf2d2019773267a7a309bc73e9d65d81978d051da2

Observation 63fb04cd-3587-417f-a374-56433a422215 · outbound

This paper cites Eyes wide shut? exploring the visual shortcomings of multimodal llms.

Beyond Sight: Towards Cognitive Alignment in LVLM via Enriched Visual Knowledge Eyes wide shut? exploring the visual shortcomings of multimodal llms

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:59:47.758643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T12:59:46.851578Z digest=sha256:9e98b70ba3b71959a9e458cfd601c3b249948645f28a02d9e9daec9af85170df

Observation 47995dcf-0171-4d26-bead-8961e82480bc · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Beyond Sight: Towards Cognitive Alignment in LVLM via Enriched Visual Knowledge Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T12:59:46.856641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:59:46.856641Z digest=sha256:1561d28be18462a4842f625ca4ce30841058d22c5b97164aaeaf5eb232bfc4db

Observation 1861502e-ac60-40ee-ba44-c646629993be · outbound

This paper cites Vary: Scaling up the vision vocabulary for large vision-language model.

Beyond Sight: Towards Cognitive Alignment in LVLM via Enriched Visual Knowledge Vary: Scaling up the vision vocabulary for large vision-language model

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:59:47.741297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T12:59:46.862040Z digest=sha256:9244db9d2fd4df94b8e179d19465acb666ed05f2a0d37675e7df72d94a70d289

Observation a31d195c-7a3b-4ff3-ac01-c4b974ac66c1 · outbound

This paper cites Weyand, A.

Beyond Sight: Towards Cognitive Alignment in LVLM via Enriched Visual Knowledge Weyand, A

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:59:47.724284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T12:59:46.867101Z digest=sha256:71c3015f51add8329d4c5a3255cc0702fe33258505436f661b3e76be70b99313

Observation 935ff633-305b-4986-9ccc-1755ed30abfe · outbound

This paper cites xgen-mm (blip-3): A family of open large multimodal models.

Beyond Sight: Towards Cognitive Alignment in LVLM via Enriched Visual Knowledge xgen-mm (blip-3): A family of open large multimodal models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-12T12:59:46.872734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:59:46.872734Z digest=sha256:f9f3f5136f23ab34da71f2c2240738e90000d6e3099840d6fd5680a223cd3f1c

Observation d5099edc-b1fe-4ed5-9b17-1d6959b6c72c · outbound

This paper cites mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models.

Beyond Sight: Towards Cognitive Alignment in LVLM via Enriched Visual Knowledge mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-12T12:59:46.878458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:59:46.878458Z digest=sha256:a7366fb43f85fd8092cada0aee9c6e31ee19772e7634177c7b6bb4b69591bb37

Observation 8f171e28-f906-4199-9611-ad7eef97c465 · outbound

This paper cites SEA: Supervised Embedding Alignment for Token-Level Visual-Textual Integration in MLLMs.

Beyond Sight: Towards Cognitive Alignment in LVLM via Enriched Visual Knowledge SEA: Supervised Embedding Alignment for Token-Level Visual-Textual Integration in MLLMs

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-12T12:59:46.883978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:59:46.883978Z digest=sha256:d4083229783636adc06f1fb139cfd2a3ae838fcbaf5b61648e71fb27cce69534

Observation 6d24949c-69b9-494a-8893-85e104b7ce4e · outbound

This paper cites Sigmoid loss for language image pre-training.

Beyond Sight: Towards Cognitive Alignment in LVLM via Enriched Visual Knowledge Sigmoid loss for language image pre-training

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-12T12:59:46.889788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:59:46.889788Z digest=sha256:7696ed5f123bd1458047abbcbce52051780809257cac41659c14bb8165678015

Observation fc142c69-688c-4b88-936e-e1e78f18748e · outbound

This paper cites InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output.

Beyond Sight: Towards Cognitive Alignment in LVLM via Enriched Visual Knowledge InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output

Reference 45

Resolution
malformed identifier
no resolver link, observed 2026-08-12T12:59:46.895927Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:59:46.895927Z digest=sha256:ba15cb8bf2350d69e29c0d4bce98e739d57116ce78e9917a2a4a79163d63498d

Observation 5d0e65b8-02c0-4adf-84c9-dc899fe83093 · outbound

This paper cites an unresolved cited work.

Beyond Sight: Towards Cognitive Alignment in LVLM via Enriched Visual Knowledge Unresolved cited work

Reference 46

Resolution
unresolved
raw_fallback, observed 2026-08-12T12:59:47.695455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T12:59:46.901670Z digest=sha256:9a3ee8a91cca735ca3b0941820c200ff6db77a1cb66d62330ca597a67f718b5f

Observation c3708779-4b38-421d-b7fa-3e306ffdafea · outbound

This paper cites an unresolved cited work.

Beyond Sight: Towards Cognitive Alignment in LVLM via Enriched Visual Knowledge Unresolved cited work

Reference 47

Resolution
unresolved
raw_fallback, observed 2026-08-12T12:59:47.677956Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T12:59:46.906835Z digest=sha256:0ff4ba9e4bc2acc871742ddd82c670ccfdaefe0cbd6b1fe42ca70e435a35e688

Observation a54bd0f3-009f-4bec-b665-3e5515fc6502 · outbound

This paper cites an unresolved cited work.

Beyond Sight: Towards Cognitive Alignment in LVLM via Enriched Visual Knowledge Unresolved cited work

Reference 48

Resolution
unresolved
raw_fallback, observed 2026-08-12T12:59:47.660662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T12:59:46.912445Z digest=sha256:530f84a11db84a0563a0513e4d8be4fa762c5b8fe82f6641012eaee8c7030370

Observation c1414642-7a9c-406d-9953-d62c308c05b6 · outbound

This paper cites an unresolved cited work.

Beyond Sight: Towards Cognitive Alignment in LVLM via Enriched Visual Knowledge Unresolved cited work

Reference 49

Resolution
unresolved
raw_fallback, observed 2026-08-12T12:59:47.643372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T12:59:46.917629Z digest=sha256:5bd9e4ffffd65d4267e62198e4b849e55375542406d1015c309a166a812d711c

Observation 30a2218d-96df-4976-8ffd-7abf937b2b19 · outbound

This paper cites an unresolved cited work.

Beyond Sight: Towards Cognitive Alignment in LVLM via Enriched Visual Knowledge Unresolved cited work

Reference 50

Resolution
unresolved
raw_fallback, observed 2026-08-12T12:59:47.626875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T12:59:46.923110Z digest=sha256:fba3f8d4b07a3ad4e9615611d44962599d7f32ab23d56dbcc8b4f37f44b41bec

Observation 6cfa80db-cba2-4f9b-8d50-dc81d775bc5e · outbound

This paper cites an unresolved cited work.

Beyond Sight: Towards Cognitive Alignment in LVLM via Enriched Visual Knowledge Unresolved cited work

Reference 51

Resolution
unresolved
raw_fallback, observed 2026-08-12T12:59:47.607932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T12:59:46.928597Z digest=sha256:6582c2d7039512410979d0e1a1e9fd5e236eae4df56a3f14cb4d82e178f37fbf

Observation ffcdece9-6c74-498c-b181-d4f09f1bf939 · outbound

This paper cites an unresolved cited work.

Beyond Sight: Towards Cognitive Alignment in LVLM via Enriched Visual Knowledge Unresolved cited work

Reference 52

Resolution
unresolved
raw_fallback, observed 2026-08-12T12:59:47.589984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T12:59:46.933580Z digest=sha256:5f80161b38e50955f0c6895e0e5a1cf5655cc7861f730dca128435c60d9515f6

Observation d7fc1fb7-f7fa-4e4b-9362-85eef8e574f6 · outbound

This paper cites an unresolved cited work.

Beyond Sight: Towards Cognitive Alignment in LVLM via Enriched Visual Knowledge Unresolved cited work

Reference 53

Resolution
unresolved
raw_fallback, observed 2026-08-12T12:59:47.569857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T12:59:46.938758Z digest=sha256:e1052c437e9edebf5873087807055d60ac3999045dc15dabb824566e038ccbd5

Observation 862a676f-81ff-436f-ac10-7ff618388ebd · outbound

This paper cites Kinderdijk Windmills“) evaluated by GPT-4o, where the answer across differ- ent models is assessed at four different levels— Strongly 3 {.

Beyond Sight: Towards Cognitive Alignment in LVLM via Enriched Visual Knowledge Kinderdijk Windmills“) evaluated by GPT-4o, where the answer across differ- ent models is assessed at four different levels— Strongly 3 {

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:59:47.550734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T12:59:46.944199Z digest=sha256:76e8fd45ef75ace8ac10373641de767192f5cc046e9513f7da10e84337059aa1

Observation 0d10a4eb-e599-4363-8e9a-7c5dc3eb2b2f · outbound

This paper cites These windmills were originally built in the 18th century to manage water levels and prevent flooding in the low-lying polder.

Beyond Sight: Towards Cognitive Alignment in LVLM via Enriched Visual Knowledge These windmills were originally built in the 18th century to manage water levels and prevent flooding in the low-lying polder

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:59:47.532107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T12:59:46.949841Z digest=sha256:1b06f82b8758c1177215f780b73fad99d93f5ea03c31b88d5f7e589e82018945

Observation 0c439667-bc9d-4c54-9d4e-029513ee4716 · outbound

This paper cites an unresolved cited work.

Beyond Sight: Towards Cognitive Alignment in LVLM via Enriched Visual Knowledge Unresolved cited work

Reference 56

Resolution
unresolved
raw_fallback, observed 2026-08-12T12:59:47.513538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T12:59:46.955076Z digest=sha256:7af0f69a3bceff6b9b012a41987463222db8a15a0ad134de834f7f6c5c11ff99

Observation 645a2dbf-57c6-4998-9d34-8dc8dd93499b · outbound

This paper cites Strongly Known,.

Beyond Sight: Towards Cognitive Alignment in LVLM via Enriched Visual Knowledge Strongly Known,

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:59:47.496456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T12:59:46.961138Z digest=sha256:646ed0d20b183ab880ceed0925de660c55560ad2ed74b8d4d2c0d21bdbe801e6

Pith citing papers

Observation 806798e9-e5f9-4bf7-bf90-b96e1a959eaa · inbound

FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL cites this paper.

FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL Beyond Sight: Towards Cognitive Alignment in LVLM via Enriched Visual Knowledge

Reference 47

Resolution
verified exact
local_arxiv, observed 2026-08-07T10:23:12.617976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T10:23:12.580211Z digest=sha256:145704c6ede9b44c59393fde8964e8dad13435012f893e5c342c83bdd76018f2