Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T12:52:30.987353Z
Paper Citation Record · LEDGER
As of 17 August 2026, this Paper Citation Record lists 75 of 75 outbound references and 4 inbound Pith citation observations for arXiv:2411.16863.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T12:52:30.987353Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-15T22:41:10.111823Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-03T20:18:57.664359Z
75 of 75 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 15193995-f07e-41ce-aa07-1f2e0931611f · outbound
Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering GPT-4 Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fc1ae605-309b-4c4b-a127-6c42907b7f5c · outbound
Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering Flamingo: a Visual Language Model for Few-Shot Learning
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 3811a72e-e950-47ad-a22d-d81e85ca10d0 · outbound
Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering Self-RAG: Learning to Retrieve, Generate, and Critique through Self-Reflection
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation a02b097d-8828-4dfb-92fc-f2de74cdb1f2 · outbound
Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c6bb8ce6-4035-486f-89f5-f3fdf01c4881 · outbound
Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering Improving Language Models by Retrieving from Trillions of Tokens
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation e13ed68d-0a65-460e-843f-a7bde0318d9d · outbound
Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering Language models are few-shot learners
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation a976634d-00b8-4173-96f1-3e3cd980c233 · outbound
Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering Tomayto, Tomahto. Beyond Token-level Answer Equivalence for Question Answering Evaluation
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d730475e-369f-4b93-8c97-3d7657e12dc7 · outbound
Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering The Revolution of Multi- modal Large Language Models: A Survey
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation f476908d-a3a9-472e-8f90-796daed1aca2 · outbound
Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering Wiki-LLaV A: Hierarchical Retrieval-Augmented Generation for Multimodal LLMs
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 2c82e968-fea9-490e-ab83-f7af28c70a93 · outbound
Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering CLAIR: Evaluating Image Captions with Large Language Models
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 0a39dd48-5d5b-4a8c-a8ca-82a9ab0fd7ec · outbound
Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering Conceptual 12M: Pushing Web-Scale Image-Text Pre-Training To Recognize Long-Tail Visual Concepts
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 349a48eb-9391-41b1-94f8-e29eda200776 · outbound
Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering Can Pre-trained Vision and Language Models Answer Visual Information-Seeking Questions? In EMNLP, 2023
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 424b804e-ee20-4b2a-9338-acb9da3293c7 · outbound
Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering Gonzalez, Ion Stoica, and Eric P
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 18f5b22e-cbd0-4adf-8d01-b47e6969d5c4 · outbound
Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering Scaling Instruction- Finetuned Language Models
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation f8133719-f5f7-4b14-8706-d58df949284d · outbound
Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering LLaVA-MORE: A Comparative Study of LLMs and Visual Backbones for Enhanced Visual Instruction Tuning
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bb9e35b7-cc96-428e-8180-1cb8a08d6140 · outbound
Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0f358bf2-de42-4505-964d-61c1b8d8531f · outbound
Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4694a8a1-da55-4535-b4b8-75eb255ba342 · outbound
Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering The Llama 3 Herd of Models
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c271dd11-ae0d-4ff5-876c-d4ff45b68277 · outbound
Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b2b5cfd5-43b7-40bf-bbb3-312b4af63d64 · outbound
Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering Data- Comp: In search of the next generation of multimodal datasets
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation ad054c70-6b26-4210-847a-87a31825449f · outbound
Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering Making the v in VQA Matter: Elevating the Role of Image Understanding in Visual Question Answer- ing
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 48f75ed0-1885-46f5-8eb1-7d1a6961810a · outbound
Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering KAT: A Knowledge Aug- mented Transformer for Vision-and-Language
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 467b33c4-640b-45e8-bd14-1ceec42634a2 · outbound
Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering Retrieval Augmented Language Model Pre- Training
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 1a18895b-4424-4585-9934-839d364cecb4 · outbound
Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering OneLLM: One Framework to Align All Modalities with Language
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 62cf2a9b-4d46-4378-8ac2-e86f40b310e1 · outbound
Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering REVEAL: Retrieval-Augmented Visual-Language Pre-Training With Multi-Source Multi- modal Knowledge Memory
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 380085ce-bd0c-442a-ad83-f1c601c217ca · outbound
Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering GQA: A New Dataset for Real-World Visual Reasoning and Compositional Question Answering
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 7ec2f541-7fd6-4e1a-81e2-89937ef0f829 · outbound
Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering Unsupervised Dense Information Retrieval with Contrastive Learning
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 988aa6a2-c74f-4989-acfd-87fed1f9ca97 · outbound
Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering Atlas: Few- shot Learning with Retrieval Augmented Language Models
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 3ecbf24f-72bd-4e8e-83b5-d54c76d4c1b6 · outbound
Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering Se- lect, Substitute, Search: A New Benchmark for Knowledge- Augmented Visual Question Answering
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 1197083e-08db-41b0-a755-b60efd128ef1 · outbound
Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering Mistral 7B
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a8385691-01c8-445a-941a-2677cf4dd2a4 · outbound
Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering Mixtral of Experts
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b6b68060-401b-41ec-9de0-6d1f24747a64 · outbound
Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering Active Retrieval Augmented Generation
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 15af0fbb-8dd7-486d-ab85-8542caaa46b3 · outbound
Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering A Diagram is Worth a Dozen Images
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 7b7eb245-60b7-44ac-8fa3-a2e67f553dc6 · outbound
Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering OBELICS: An Open Web-Scale Filtered Dataset of Inter- leaved Image-Text Documents
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 48228310-b2bd-4ffb-a557-8e133048519b · outbound
Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering What matters when building vision-language models?
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f0df665a-15a1-4fb6-8892-8f7035493f72 · outbound
Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering ViQuAE, a dataset for knowledge-based visual question answering about named entities
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation f32c6df3-1330-4cb4-b1bf-0baee84217fe · outbound
Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering Cross- modal Retrieval for Knowledge-based Visual Question An- swering
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 6af941cb-0f15-4633-ab16-83bbbc988216 · outbound
Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4fbfe0a5-fc83-4d63-adf2-d29e7d06adb0 · outbound
Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering LLaVA-OneVision: Easy Visual Task Transfer
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e57701ce-6595-4167-b19a-4f16073a46bf · outbound
Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering BLIP- 2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation d1680a28-1b1f-427f-b04b-b59ffb39465a · outbound
Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering Evaluating Object Hallucination in Large Vision-Language Models
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 817344e4-8107-47a1-a385-ba67d2f9600b · outbound
Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering Fine-grained Late-interaction Multi-modal Retrieval for Retrieval Augmented Visual Question Answer- ing
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation c958afb6-acc9-4701-b9e2-9bcea11e72ae · outbound
Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering PreFLMR: Scaling Up Fine-Grained Late-Interaction Multi- modal Retrievers
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 19d3b4f5-a0a7-46a5-9a46-b149035e295d · outbound
Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering REVIVE: Regional Visual Rep- resentation Matters in Knowledge-Based Visual Question Answering
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 420a1918-7ba8-4da0-9c40-93a66b5a5cbe · outbound
Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering Visual Instruction Tuning
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation ca5f192d-0623-49be-910a-b64e9d12158c · outbound
Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering Improved Baselines with Visual Instruction Tuning
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation a5707012-71a3-481c-9bc5-2976e54a613d · outbound
Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering MMBench: Is Your Multi-modal Model an All-around Player?
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a51a05f5-e4f0-4403-948f-bbb4bd55400d · outbound
Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question Answering
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation b04fdc93-75e9-4cf6-9c77-048753ca6730 · outbound
Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering Ok-VQA: A Visual Question Answering Benchmark Requiring External Knowledge
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 27ef76d9-b44b-46c6-9511-e5b546c06af8 · outbound
Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering MM1: Methods, Analy- sis & Insights from Multimodal LLM Pre-training
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 98a48acc-3c2c-46a0-8600-9995ff825a27 · outbound
Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering Encyclopedic VQA: Visual Questions About Detailed Properties of Fine-Grained Categories
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 31935bd8-1fd6-4e7c-bb97-7bfc22cd9ec7 · outbound
Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering PlotQA: Reasoning over Scientific Plots
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 257e8f23-de4a-4de9-b5e5-68a594c6303a · outbound
Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering Revisiting Image Cap- tioning Training Paradigm via Direct CLIP-based Optimiza- tion
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 155570d5-f38f-4430-840d-a3a0a51a6311 · outbound
Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering Training Language Models to Follow Instructions with Human Feedback
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation e5817235-ee81-48fb-9eb4-e46b09e94648 · outbound
Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering RoRA-VLM: Robust Retrieval-Augmented Vision Language Models
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f4ca4721-3864-4d6d-bb24-2540961054c4 · outbound
Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering Learning Transferable Visual Models from Natural Language Supervi- sion
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 3c97973b-06d1-4bfd-8cc8-61ce028e3368 · outbound
Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 703ed348-9f2d-4986-b079-dc40e27e3f5d · outbound
Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering In- Context Retrieval-Augmented Language Models
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation ce2f2d31-a14c-41ac-a703-1e7baf963715 · outbound
Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering A-OKVQA: A Benchmark for Visual Question Answering Using World Knowledge
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation a7541de2-06fb-46d7-a623-70e4be65e50b · outbound
Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering KVQA: Knowledge-aware Visual Question Answering
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation d825cc38-08a9-49b3-be16-bfa5fa793f83 · outbound
Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering Towards VQA Models That Can Read
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation b9ba3b35-fc43-4983-8e92-e5a8a5022844 · outbound
Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering WIT: Wikipedia-based Image Text Dataset for Multimodal Multilingual Machine Learning
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 604dbd24-601d-4acc-8380-8d19cf97cc55 · outbound
Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 446863de-a2e6-4ac4-8d64-66041767d9d4 · outbound
Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering Emu: Generative Pretraining in Multimodality
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 3d914263-29fa-4526-bd81-7bbf64eb7de8 · outbound
Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering Stanford Alpaca: An Instruction-Following LLaMA Model, 2023
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation b09f0594-c9af-487e-b06c-6714b6712f01 · outbound
Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering LLaMA: Open and Efficient Foundation Language Models
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ed0a2a25-31f6-433b-91d6-781c0281c5b8 · outbound
Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering InstructRetro: Instruction Tuning post Retrieval-Augmented Pretraining
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 0e81e5b5-91de-415a-acc8-ba3e753f4aa6 · outbound
Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering UniIR: Training and Benchmarking Universal Multimodal Information Retrievers
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 493ee015-4639-4ca9-b37c-63b8946309e5 · outbound
Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering CLAIR-A: Leveraging Large Language Models to Judge Audio Captions
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 56af160e-9fa7-4545-810e-5338275d0ded · outbound
Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering Grounding Language Models for Visual Entity Recognition
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 39901545-cd60-41e3-ba81-fc266361067d · outbound
Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering EchoSight: Advancing Visual-Language Models with Wiki Knowledge
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2226b6c7-5663-403e-b31a-0275c673d69c · outbound
Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering mPLUG- Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation b59bc50d-3ed5-446c-9c78-48ade2622f41 · outbound
Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering RankRAG: Unifying Context Ranking with Retrieval-Augmented Generation in LLMs
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 255f2df9-65e1-405a-ba53-b125d817d9b6 · outbound
Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering Give a short answer
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 39fc22b7-cffb-4284-a342-eddaa3a21928 · outbound
Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering Unresolved cited work
Reference 2024
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation b7ea84fc-69d5-4fb1-baca-2570f8eedaaa · inbound
OMGM: Orchestrate Multiple Granularities and Modalities for Efficient Multimodal Retrieval Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8109059b-35dd-4671-8911-22da5803b3ed · inbound
Towards General Continuous Memory for Vision-Language Models Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation da91c099-b84d-4551-8bd6-523da596ae4e · inbound
ExpStar: Towards Automatic Commentary Generation for Multi-discipline Scientific Experiments Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4882e389-ddcf-4108-8eca-840a641f684d · inbound
MathVis-Fine: Aligning Visual Supervision with Necessity via Progressive Dependency-Guided Training for Multimodal Mathematical Reasoning Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.