Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T12:59:46.961138Z
Paper Citation Record · LEDGER
As of 15 August 2026, this Paper Citation Record lists 57 of 57 outbound references and 1 inbound Pith citation observation for arXiv:2411.16824.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T12:59:46.961138Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T10:23:12.580211Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-07T10:23:12.613702Z
57 of 57 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation a0695d83-0a85-4631-9a85-397e1c593e00 · outbound
Beyond Sight: Towards Cognitive Alignment in LVLM via Enriched Visual Knowledge Llama 3 model card
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 08c112a0-3e16-4516-99e1-15bd5f9f681b · outbound
Beyond Sight: Towards Cognitive Alignment in LVLM via Enriched Visual Knowledge Flamingo: a visual language model for few-shot learning
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bfcb16bb-701d-4870-a262-d996e2aba22a · outbound
Beyond Sight: Towards Cognitive Alignment in LVLM via Enriched Visual Knowledge Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eaf489e2-279c-4785-a40d-7a05b0ce8825 · outbound
Beyond Sight: Towards Cognitive Alignment in LVLM via Enriched Visual Knowledge Improving fine-grained understanding in image-text pre-training
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a056aa9-3c99-4753-89d4-59e941578d09 · outbound
Beyond Sight: Towards Cognitive Alignment in LVLM via Enriched Visual Knowledge Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f343a77a-991f-4c2d-ae99-e750c5b2b9b7 · outbound
Beyond Sight: Towards Cognitive Alignment in LVLM via Enriched Visual Knowledge InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fb9168b0-f1d2-4eda-83df-06b59c5464d1 · outbound
Beyond Sight: Towards Cognitive Alignment in LVLM via Enriched Visual Knowledge InternLM-XComposer2-4KHD: A Pioneering Large Vision-Language Model Handling Resolutions from 336 Pixels to 4K HD
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e8910bbd-b644-42c1-985d-6c6d9f941408 · outbound
Beyond Sight: Towards Cognitive Alignment in LVLM via Enriched Visual Knowledge Does Fine-Tuning LLMs on New Knowledge Encourage Hallucinations?
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7bafba62-0bf5-4efd-a4ca-2a7e7a3dae25 · outbound
Beyond Sight: Towards Cognitive Alignment in LVLM via Enriched Visual Knowledge Bard, 2023
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6f67d8a9-a55a-40d9-a70b-aeabf8ff9c47 · outbound
Beyond Sight: Towards Cognitive Alignment in LVLM via Enriched Visual Knowledge Making the V in VQA matter: El- evating the role of image understanding in visual question answering
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 5b567098-450f-4c77-a295-d70ed847689a · outbound
Beyond Sight: Towards Cognitive Alignment in LVLM via Enriched Visual Knowledge GQA: A new dataset for real-world visual reasoning and compositional question answering
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation a6228346-387d-4473-a46f-d252efea1e45 · outbound
Beyond Sight: Towards Cognitive Alignment in LVLM via Enriched Visual Knowledge BRAVE: Broadening the visual encoding of vision-language models
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7c77ee76-4c8b-407a-9d52-51a893626f05 · outbound
Beyond Sight: Towards Cognitive Alignment in LVLM via Enriched Visual Knowledge Referitgame: Referring to objects in pho- tographs of natural scenes
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 40b7d9a5-72b7-41a7-8620-2be600bc6ff9 · outbound
Beyond Sight: Towards Cognitive Alignment in LVLM via Enriched Visual Knowledge Visual genome: Connecting language and vision using crowdsourced dense image annotations
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation af80f546-e2ee-458d-95ac-5b6f4a554bf2 · outbound
Beyond Sight: Towards Cognitive Alignment in LVLM via Enriched Visual Knowledge What matters when building vision-language models?
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 206f8b15-97c3-4580-b4fa-10851a71432e · outbound
Beyond Sight: Towards Cognitive Alignment in LVLM via Enriched Visual Knowledge BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a31b0483-7240-4424-97f5-13afc61b19f1 · outbound
Beyond Sight: Towards Cognitive Alignment in LVLM via Enriched Visual Knowledge Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d7cabda7-85ed-4be5-9778-0f642e98ae99 · outbound
Beyond Sight: Towards Cognitive Alignment in LVLM via Enriched Visual Knowledge SPHINX-X: Scaling Data and Parameters for a Family of Multi-modal Large Language Models
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f20cbc10-451d-4e7f-b18f-41c17b56841b · outbound
Beyond Sight: Towards Cognitive Alignment in LVLM via Enriched Visual Knowledge Improved Baselines with Visual Instruction Tuning
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2ceadc13-e658-4322-8c20-7c5dd733449c · outbound
Beyond Sight: Towards Cognitive Alignment in LVLM via Enriched Visual Knowledge Visual instruction tuning
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 90a69544-958d-486e-a519-5d9d6b57ce32 · outbound
Beyond Sight: Towards Cognitive Alignment in LVLM via Enriched Visual Knowledge Decoupled weight decay regularization
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 785dbdae-3e40-491d-aced-212468d13fc8 · outbound
Beyond Sight: Towards Cognitive Alignment in LVLM via Enriched Visual Knowledge DeepSeek-VL: Towards Real-World Vision-Language Understanding
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b930b0c0-a931-402a-a0e0-ecbb4ee55ca2 · outbound
Beyond Sight: Towards Cognitive Alignment in LVLM via Enriched Visual Knowledge Generation and comprehension of unambiguous object descriptions
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1df21175-905e-4737-94b2-17a3c7600c80 · outbound
Beyond Sight: Towards Cognitive Alignment in LVLM via Enriched Visual Knowledge OK-VQA: A visual question answering benchmark requiring external knowledge
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 3f028d67-e504-41d0-889c-0bbf12920656 · outbound
Beyond Sight: Towards Cognitive Alignment in LVLM via Enriched Visual Knowledge OCR-VQA: Visual question answer- ing by reading text in images
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 27186e34-2e09-444b-bf47-f7f97774a9df · outbound
Beyond Sight: Towards Cognitive Alignment in LVLM via Enriched Visual Knowledge Gpt-4 technical report, 2023
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation dd7db9c0-3ba7-42f8-922a-a90bf145094b · outbound
Beyond Sight: Towards Cognitive Alignment in LVLM via Enriched Visual Knowledge GPT-4o System Card, 2024
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 664e6cba-6119-4a20-8a07-33b712634612 · outbound
Beyond Sight: Towards Cognitive Alignment in LVLM via Enriched Visual Knowledge Instruction Tuning with GPT-4
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2f5af47a-8438-41c3-a1f1-5726e5b82249 · outbound
Beyond Sight: Towards Cognitive Alignment in LVLM via Enriched Visual Knowledge Learning transferable visual models from natural language supervision, 2021
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation f2379740-2995-44d1-8a84-7211cdd51519 · outbound
Beyond Sight: Towards Cognitive Alignment in LVLM via Enriched Visual Knowledge A-OKVQA: A benchmark for visual question answering using world knowl- edge
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation a97eb258-fb28-49d7-893f-53b94bb17d46 · outbound
Beyond Sight: Towards Cognitive Alignment in LVLM via Enriched Visual Knowledge Conceptual captions: A cleaned, hypernymed, im- age alt-text dataset for automatic image captioning
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 57d375b8-60f2-45ec-b788-291bb1f1c070 · outbound
Beyond Sight: Towards Cognitive Alignment in LVLM via Enriched Visual Knowledge When do we not need larger vision models? In European Conference on Computer Vision, pages 444–462
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation f76eff6c-0d61-4de5-873a-f521408de1a1 · outbound
Beyond Sight: Towards Cognitive Alignment in LVLM via Enriched Visual Knowledge Textcaps: a dataset for image captioning with reading comprehension
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 98f624ac-cd17-420f-841f-6603d7b795cb · outbound
Beyond Sight: Towards Cognitive Alignment in LVLM via Enriched Visual Knowledge EVA-CLIP: Improved Training Techniques for CLIP at Scale
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b915ea51-373b-4822-965c-37ef940d6ae0 · outbound
Beyond Sight: Towards Cognitive Alignment in LVLM via Enriched Visual Knowledge When are lemons purple? the concept association bias of vision-language models
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 92ac9bfa-b275-4876-a148-4bcdad1dad42 · outbound
Beyond Sight: Towards Cognitive Alignment in LVLM via Enriched Visual Knowledge Qwen2.5: A party of foundation models, 2024
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 63fb04cd-3587-417f-a374-56433a422215 · outbound
Beyond Sight: Towards Cognitive Alignment in LVLM via Enriched Visual Knowledge Eyes wide shut? exploring the visual shortcomings of multimodal llms
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 47995dcf-0171-4d26-bead-8961e82480bc · outbound
Beyond Sight: Towards Cognitive Alignment in LVLM via Enriched Visual Knowledge Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1861502e-ac60-40ee-ba44-c646629993be · outbound
Beyond Sight: Towards Cognitive Alignment in LVLM via Enriched Visual Knowledge Vary: Scaling up the vision vocabulary for large vision-language model
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation a31d195c-7a3b-4ff3-ac01-c4b974ac66c1 · outbound
Beyond Sight: Towards Cognitive Alignment in LVLM via Enriched Visual Knowledge Weyand, A
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 935ff633-305b-4986-9ccc-1755ed30abfe · outbound
Beyond Sight: Towards Cognitive Alignment in LVLM via Enriched Visual Knowledge xgen-mm (blip-3): A family of open large multimodal models
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d5099edc-b1fe-4ed5-9b17-1d6959b6c72c · outbound
Beyond Sight: Towards Cognitive Alignment in LVLM via Enriched Visual Knowledge mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8f171e28-f906-4199-9611-ad7eef97c465 · outbound
Beyond Sight: Towards Cognitive Alignment in LVLM via Enriched Visual Knowledge SEA: Supervised Embedding Alignment for Token-Level Visual-Textual Integration in MLLMs
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6d24949c-69b9-494a-8893-85e104b7ce4e · outbound
Beyond Sight: Towards Cognitive Alignment in LVLM via Enriched Visual Knowledge Sigmoid loss for language image pre-training
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fc142c69-688c-4b88-936e-e1e78f18748e · outbound
Beyond Sight: Towards Cognitive Alignment in LVLM via Enriched Visual Knowledge InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d0e65b8-02c0-4adf-84c9-dc899fe83093 · outbound
Beyond Sight: Towards Cognitive Alignment in LVLM via Enriched Visual Knowledge Unresolved cited work
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation c3708779-4b38-421d-b7fa-3e306ffdafea · outbound
Beyond Sight: Towards Cognitive Alignment in LVLM via Enriched Visual Knowledge Unresolved cited work
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation a54bd0f3-009f-4bec-b665-3e5515fc6502 · outbound
Beyond Sight: Towards Cognitive Alignment in LVLM via Enriched Visual Knowledge Unresolved cited work
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation c1414642-7a9c-406d-9953-d62c308c05b6 · outbound
Beyond Sight: Towards Cognitive Alignment in LVLM via Enriched Visual Knowledge Unresolved cited work
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 30a2218d-96df-4976-8ffd-7abf937b2b19 · outbound
Beyond Sight: Towards Cognitive Alignment in LVLM via Enriched Visual Knowledge Unresolved cited work
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 6cfa80db-cba2-4f9b-8d50-dc81d775bc5e · outbound
Beyond Sight: Towards Cognitive Alignment in LVLM via Enriched Visual Knowledge Unresolved cited work
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation ffcdece9-6c74-498c-b181-d4f09f1bf939 · outbound
Beyond Sight: Towards Cognitive Alignment in LVLM via Enriched Visual Knowledge Unresolved cited work
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation d7fc1fb7-f7fa-4e4b-9362-85eef8e574f6 · outbound
Beyond Sight: Towards Cognitive Alignment in LVLM via Enriched Visual Knowledge Unresolved cited work
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 862a676f-81ff-436f-ac10-7ff618388ebd · outbound
Beyond Sight: Towards Cognitive Alignment in LVLM via Enriched Visual Knowledge Kinderdijk Windmills“) evaluated by GPT-4o, where the answer across differ- ent models is assessed at four different levels— Strongly 3 {
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 0d10a4eb-e599-4363-8e9a-7c5dc3eb2b2f · outbound
Beyond Sight: Towards Cognitive Alignment in LVLM via Enriched Visual Knowledge These windmills were originally built in the 18th century to manage water levels and prevent flooding in the low-lying polder
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 0c439667-bc9d-4c54-9d4e-029513ee4716 · outbound
Beyond Sight: Towards Cognitive Alignment in LVLM via Enriched Visual Knowledge Unresolved cited work
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 645a2dbf-57c6-4998-9d34-8dc8dd93499b · outbound
Beyond Sight: Towards Cognitive Alignment in LVLM via Enriched Visual Knowledge Strongly Known,
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 806798e9-e5f9-4bf7-bf90-b96e1a959eaa · inbound
FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL Beyond Sight: Towards Cognitive Alignment in LVLM via Enriched Visual Knowledge
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.