Pith. sign in

Paper Citation Record · LEDGER

Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering

As of 17 August 2026, this Paper Citation Record lists 75 of 75 outbound references and 4 inbound Pith citation observations for arXiv:2411.16863.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.16863 v2

Coverage vector

measured 75 of 75 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T12:52:30.987353Z

measured 79 of 79 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T22:41:10.111823Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T20:18:57.664359Z

Reference resolution

75 of 75 outbound references displayed

  • verified exact1
  • verified fuzzy52
  • unresolved20
  • parse uncertain1
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 15193995-f07e-41ce-aa07-1f2e0931611f · outbound

This paper cites GPT-4 Technical Report.

Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T12:52:30.636989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:52:30.636989Z digest=sha256:7f08af637514a0cb963a761ff167df7e4f8d98ce297d5cb771b71b127c75145b

Observation fc1ae605-309b-4c4b-a127-6c42907b7f5c · outbound

This paper cites Flamingo: a Visual Language Model for Few-Shot Learning.

Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering Flamingo: a Visual Language Model for Few-Shot Learning

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:52:32.122624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T12:52:30.642515Z digest=sha256:e87a37f5a2afdbb56c54fa0b71dafd1d96a9c23f6d43a2d396d5987dc08f52f1

Observation 3811a72e-e950-47ad-a22d-d81e85ca10d0 · outbound

This paper cites Self-RAG: Learning to Retrieve, Generate, and Critique through Self-Reflection.

Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering Self-RAG: Learning to Retrieve, Generate, and Critique through Self-Reflection

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:52:32.107373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T12:52:30.647473Z digest=sha256:05452bbb0eaaa770382a34375faa13fa452a01456244838384e4bea808a153a5

Observation a02b097d-8828-4dfb-92fc-f2de74cdb1f2 · outbound

This paper cites OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models.

Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T12:52:30.652319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:52:30.652319Z digest=sha256:c989f35a5619dedf75b7b17875003f499b6d1f3480d60744bb29e19f3a1da749

Observation c6bb8ce6-4035-486f-89f5-f3fdf01c4881 · outbound

This paper cites Improving Language Models by Retrieving from Trillions of Tokens.

Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering Improving Language Models by Retrieving from Trillions of Tokens

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:52:32.092469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T12:52:30.658292Z digest=sha256:2fca3735defefb7c568a5edf6a4bde836b02e78d262688c16b5c65e7a85b8a20

Observation e13ed68d-0a65-460e-843f-a7bde0318d9d · outbound

This paper cites Language models are few-shot learners.

Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering Language models are few-shot learners

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:52:32.078053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T12:52:30.663639Z digest=sha256:123a5aab5c816e52c8f925b904d40119120c6ee4b61adbac3b1a36464ba8a3f5

Observation a976634d-00b8-4173-96f1-3e3cd980c233 · outbound

This paper cites Tomayto, Tomahto. Beyond Token-level Answer Equivalence for Question Answering Evaluation.

Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering Tomayto, Tomahto. Beyond Token-level Answer Equivalence for Question Answering Evaluation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T12:52:30.668611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:52:30.668611Z digest=sha256:519816888625c59d8582c369ecf0fe48012934105ec2b59bdeea9b90cefe4cd0

Observation d730475e-369f-4b93-8c97-3d7657e12dc7 · outbound

This paper cites The Revolution of Multi- modal Large Language Models: A Survey.

Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering The Revolution of Multi- modal Large Language Models: A Survey

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:52:32.063688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T12:52:30.673955Z digest=sha256:d014a05d2a12303773151d49b63cdd3d362ecbc4ef8cbc7f72c86a5479478cca

Observation f476908d-a3a9-472e-8f90-796daed1aca2 · outbound

This paper cites Wiki-LLaV A: Hierarchical Retrieval-Augmented Generation for Multimodal LLMs.

Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering Wiki-LLaV A: Hierarchical Retrieval-Augmented Generation for Multimodal LLMs

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:52:32.048863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T12:52:30.678609Z digest=sha256:a6d741197f8f8086bb08eb4e3c928503168a2f7059418b31e70d45caf63f68d9

Observation 2c82e968-fea9-490e-ab83-f7af28c70a93 · outbound

This paper cites CLAIR: Evaluating Image Captions with Large Language Models.

Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering CLAIR: Evaluating Image Captions with Large Language Models

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:52:32.033761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T12:52:30.683205Z digest=sha256:ea26562186fa64215243964761796de4e99e19535d5ae0c809546fca5d5eaf54

Observation 0a39dd48-5d5b-4a8c-a8ca-82a9ab0fd7ec · outbound

This paper cites Conceptual 12M: Pushing Web-Scale Image-Text Pre-Training To Recognize Long-Tail Visual Concepts.

Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering Conceptual 12M: Pushing Web-Scale Image-Text Pre-Training To Recognize Long-Tail Visual Concepts

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:52:32.018859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T12:52:30.688031Z digest=sha256:18a3b5cd28f9ceb99cc78708229b482b2793e4f28ef5fbcead8b317ccb4a8c6b

Observation 349a48eb-9391-41b1-94f8-e29eda200776 · outbound

This paper cites Can Pre-trained Vision and Language Models Answer Visual Information-Seeking Questions? In EMNLP, 2023.

Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering Can Pre-trained Vision and Language Models Answer Visual Information-Seeking Questions? In EMNLP, 2023

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:52:32.004116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T12:52:30.692380Z digest=sha256:d613507c37df9ed57b678a063fbbf0ea276ed6afac9c58e9d82deeccffd2db02

Observation 424b804e-ee20-4b2a-9338-acb9da3293c7 · outbound

This paper cites Gonzalez, Ion Stoica, and Eric P.

Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering Gonzalez, Ion Stoica, and Eric P

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:52:31.989168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T12:52:30.696826Z digest=sha256:4394d3c933196f7e6d8c0539eb5ae832d1ceec680986e69d3dae7c886ad458c9

Observation 18f5b22e-cbd0-4adf-8d01-b47e6969d5c4 · outbound

This paper cites Scaling Instruction- Finetuned Language Models.

Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering Scaling Instruction- Finetuned Language Models

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:52:31.975256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T12:52:30.701200Z digest=sha256:2b1fea61173440a552e1be0d709df9260c061abe15ec856a94c1cd01a6bf8745

Observation f8133719-f5f7-4b14-8706-d58df949284d · outbound

This paper cites LLaVA-MORE: A Comparative Study of LLMs and Visual Backbones for Enhanced Visual Instruction Tuning.

Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering LLaVA-MORE: A Comparative Study of LLMs and Visual Backbones for Enhanced Visual Instruction Tuning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T12:52:30.705368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:52:30.705368Z digest=sha256:34eb78600b95d5bb7c6073e6293f4b2a6c9b038d4fdfc494be6f86c3d7e729a6

Observation bb9e35b7-cc96-428e-8180-1cb8a08d6140 · outbound

This paper cites InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning.

Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T12:52:30.710100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:52:30.710100Z digest=sha256:d1c826e7f21b0b5ccbfbc3283b8b3cd40ba6d462567fff067c34c52b8b5a26dd

Observation 0f358bf2-de42-4505-964d-61c1b8d8531f · outbound

This paper cites Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models.

Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T12:52:30.714869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:52:30.714869Z digest=sha256:2337ff9f26f1247dcd33070b5dc58cdbb4625e1abb529ff51817456b9ba19007

Observation 4694a8a1-da55-4535-b4b8-75eb255ba342 · outbound

This paper cites The Llama 3 Herd of Models.

Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering The Llama 3 Herd of Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T12:52:30.719562Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:52:30.719562Z digest=sha256:30fb41fa401ba84b467ab86735dd379d352de9a215cf9fa49a12aa56327e3b0c

Observation c271dd11-ae0d-4ff5-876c-d4ff45b68277 · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T12:52:30.724605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:52:30.724605Z digest=sha256:48eebc2acba2e2227fa77f46c78c40958bfd807a034746c8194747875735660e

Observation b2b5cfd5-43b7-40bf-bbb3-312b4af63d64 · outbound

This paper cites Data- Comp: In search of the next generation of multimodal datasets.

Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering Data- Comp: In search of the next generation of multimodal datasets

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:52:31.960666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T12:52:30.729225Z digest=sha256:cfd1484f2ba74688fae1364806e5bf70c3f56847c3a7a99078c3c7f5baf3c6a6

Observation ad054c70-6b26-4210-847a-87a31825449f · outbound

This paper cites Making the v in VQA Matter: Elevating the Role of Image Understanding in Visual Question Answer- ing.

Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering Making the v in VQA Matter: Elevating the Role of Image Understanding in Visual Question Answer- ing

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:52:31.944607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T12:52:30.733644Z digest=sha256:dac981a637b1260a271ef618c1f74c8c06aae6096f27e7c14c1fbb97cfa7435b

Observation 48f75ed0-1885-46f5-8eb1-7d1a6961810a · outbound

This paper cites KAT: A Knowledge Aug- mented Transformer for Vision-and-Language.

Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering KAT: A Knowledge Aug- mented Transformer for Vision-and-Language

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:52:31.929624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T12:52:30.737977Z digest=sha256:823353587ab14088fede4d49d09a66446df956902fbe7cf925f3e3ccbbe98a5e

Observation 467b33c4-640b-45e8-bd14-1ceec42634a2 · outbound

This paper cites Retrieval Augmented Language Model Pre- Training.

Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering Retrieval Augmented Language Model Pre- Training

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:52:31.915584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T12:52:30.742323Z digest=sha256:91554d7bb5cce891b0f4687330bb3983a73868f47b6f58227ad770332f9f9b8a

Observation 1a18895b-4424-4585-9934-839d364cecb4 · outbound

This paper cites OneLLM: One Framework to Align All Modalities with Language.

Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering OneLLM: One Framework to Align All Modalities with Language

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:52:31.901483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T12:52:30.746903Z digest=sha256:090ead85b82f04ab398364432a4e229a3756e701e42be55614ddb2bbb297866b

Observation 62cf2a9b-4d46-4378-8ac2-e86f40b310e1 · outbound

This paper cites REVEAL: Retrieval-Augmented Visual-Language Pre-Training With Multi-Source Multi- modal Knowledge Memory.

Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering REVEAL: Retrieval-Augmented Visual-Language Pre-Training With Multi-Source Multi- modal Knowledge Memory

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:52:31.887720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T12:52:30.751442Z digest=sha256:f2e150944648b219c226e6a1202f3f17f49d9ded06a47317b51afdc0eae5a9e8

Observation 380085ce-bd0c-442a-ad83-f1c601c217ca · outbound

This paper cites GQA: A New Dataset for Real-World Visual Reasoning and Compositional Question Answering.

Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering GQA: A New Dataset for Real-World Visual Reasoning and Compositional Question Answering

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:52:31.873422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T12:52:30.755733Z digest=sha256:5c0fbdbbb77d3c31739213350b0929906bf84302f0b3a35d5bbf529f80884303

Observation 7ec2f541-7fd6-4e1a-81e2-89937ef0f829 · outbound

This paper cites Unsupervised Dense Information Retrieval with Contrastive Learning.

Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering Unsupervised Dense Information Retrieval with Contrastive Learning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T12:52:30.760931Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:52:30.760931Z digest=sha256:9d22ed7365ecdc689508d04c2d5e7bf726400241668e6286d4f25cd7ceb1d7cc

Observation 988aa6a2-c74f-4989-acfd-87fed1f9ca97 · outbound

This paper cites Atlas: Few- shot Learning with Retrieval Augmented Language Models.

Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering Atlas: Few- shot Learning with Retrieval Augmented Language Models

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:52:31.859421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T12:52:30.765665Z digest=sha256:b8369c47b6cd908e8a01d327b6e574022372d462fb0dcf98e42ab79aa871be02

Observation 3ecbf24f-72bd-4e8e-83b5-d54c76d4c1b6 · outbound

This paper cites Se- lect, Substitute, Search: A New Benchmark for Knowledge- Augmented Visual Question Answering.

Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering Se- lect, Substitute, Search: A New Benchmark for Knowledge- Augmented Visual Question Answering

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:52:31.844961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T12:52:30.770143Z digest=sha256:e3703d2337c2f0a8c2d7d98965bcacfc1aa0f0fb3cc9bbbd13d648a26d1e6281

Observation 1197083e-08db-41b0-a755-b60efd128ef1 · outbound

This paper cites Mistral 7B.

Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering Mistral 7B

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T12:52:30.774913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:52:30.774913Z digest=sha256:cac444624acefa52193fff45f7997ef8225f98c3ee25450398136d908358ba9f

Observation a8385691-01c8-445a-941a-2677cf4dd2a4 · outbound

This paper cites Mixtral of Experts.

Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering Mixtral of Experts

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T12:52:30.779780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:52:30.779780Z digest=sha256:267ff1d5a40b2f7e8afb5d76d41e68b54835421e5eaa241d74ac716227ec3e97

Observation b6b68060-401b-41ec-9de0-6d1f24747a64 · outbound

This paper cites Active Retrieval Augmented Generation.

Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering Active Retrieval Augmented Generation

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:52:31.824941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T12:52:30.784523Z digest=sha256:686995748c43de3492603bd9236dac4493d73f092b259f31b8db053b521ff737

Observation 15af0fbb-8dd7-486d-ab85-8542caaa46b3 · outbound

This paper cites A Diagram is Worth a Dozen Images.

Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering A Diagram is Worth a Dozen Images

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:52:31.810453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T12:52:30.788927Z digest=sha256:cce097aa4f0dfa5471e14754ac71d81933f1dffd68780ac4f876b969e27c5afe

Observation 7b7eb245-60b7-44ac-8fa3-a2e67f553dc6 · outbound

This paper cites OBELICS: An Open Web-Scale Filtered Dataset of Inter- leaved Image-Text Documents.

Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering OBELICS: An Open Web-Scale Filtered Dataset of Inter- leaved Image-Text Documents

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:52:31.795675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T12:52:30.793333Z digest=sha256:f5b32107a705cf1a4df42d3588b4fa479b9ec9e3917fa8c879164dda5842579a

Observation 48228310-b2bd-4ffb-a557-8e133048519b · outbound

This paper cites What matters when building vision-language models?.

Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering What matters when building vision-language models?

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-12T12:52:30.798029Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:52:30.798029Z digest=sha256:3cf214e2126d8977d13ca493dec093c12612b1abe0894e4bd8f61852fccce260

Observation f0df665a-15a1-4fb6-8892-8f7035493f72 · outbound

This paper cites ViQuAE, a dataset for knowledge-based visual question answering about named entities.

Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering ViQuAE, a dataset for knowledge-based visual question answering about named entities

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:52:31.781277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T12:52:30.802885Z digest=sha256:da050a4010731093709ecc45a9370dbe98b7e26446f54190bf339b160f2774ad

Observation f32c6df3-1330-4cb4-b1bf-0baee84217fe · outbound

This paper cites Cross- modal Retrieval for Knowledge-based Visual Question An- swering.

Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering Cross- modal Retrieval for Knowledge-based Visual Question An- swering

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:52:31.766226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T12:52:30.808466Z digest=sha256:e02e1e17f77abe24e8d97f8314b9206f98f746ce2b959fff86e6e7216e755a07

Observation 6af941cb-0f15-4633-ab16-83bbbc988216 · outbound

This paper cites SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension.

Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T12:52:30.812696Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:52:30.812696Z digest=sha256:e2d9e4725800c7c43da353cf4b78a1fca35b86eb009ad788994198f53ff874d7

Observation 4fbfe0a5-fc83-4d63-adf2-d29e7d06adb0 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering LLaVA-OneVision: Easy Visual Task Transfer

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-12T12:52:30.817373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:52:30.817373Z digest=sha256:38373e4bf74c700199980817b7b89ae818bace5901b34fa5bd6dc690597207a8

Observation e57701ce-6595-4167-b19a-4f16073a46bf · outbound

This paper cites BLIP- 2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models.

Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering BLIP- 2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:52:31.751830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T12:52:30.822147Z digest=sha256:8b7f3d85656954f2b5e8e116e8b7029fb12c06bf22070f6b9c11aed988994654

Observation d1680a28-1b1f-427f-b04b-b59ffb39465a · outbound

This paper cites Evaluating Object Hallucination in Large Vision-Language Models.

Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering Evaluating Object Hallucination in Large Vision-Language Models

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:52:31.737968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T12:52:30.827570Z digest=sha256:96abd3c7cca2f26020d237a2df10c5f9cb07f1b4a4234133a6a58709a26338f8

Observation 817344e4-8107-47a1-a385-ba67d2f9600b · outbound

This paper cites Fine-grained Late-interaction Multi-modal Retrieval for Retrieval Augmented Visual Question Answer- ing.

Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering Fine-grained Late-interaction Multi-modal Retrieval for Retrieval Augmented Visual Question Answer- ing

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:52:31.724309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T12:52:30.832476Z digest=sha256:26fd84f51df09349afec0b16fe49a42d22246a9089a95597b8989bce6f37c6ba

Observation c958afb6-acc9-4701-b9e2-9bcea11e72ae · outbound

This paper cites PreFLMR: Scaling Up Fine-Grained Late-Interaction Multi- modal Retrievers.

Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering PreFLMR: Scaling Up Fine-Grained Late-Interaction Multi- modal Retrievers

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:52:31.709865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T12:52:30.837104Z digest=sha256:7d79aac3a7c794b20ca30f3f69070c6938526fe3015be1561e2e94191af02887

Observation 19d3b4f5-a0a7-46a5-9a46-b149035e295d · outbound

This paper cites REVIVE: Regional Visual Rep- resentation Matters in Knowledge-Based Visual Question Answering.

Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering REVIVE: Regional Visual Rep- resentation Matters in Knowledge-Based Visual Question Answering

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:52:31.695207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T12:52:30.841544Z digest=sha256:521a91d19ce329a52302f68ab44054512380cc8f3d8efb946c1b0248b523816c

Observation 420a1918-7ba8-4da0-9c40-93a66b5a5cbe · outbound

This paper cites Visual Instruction Tuning.

Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering Visual Instruction Tuning

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:52:31.681243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T12:52:30.845959Z digest=sha256:162d63674c4aef6b9e5457c722efe28393a3f367f0607ba69163779167fa3a64

Observation ca5f192d-0623-49be-910a-b64e9d12158c · outbound

This paper cites Improved Baselines with Visual Instruction Tuning.

Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering Improved Baselines with Visual Instruction Tuning

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:52:31.666849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T12:52:30.850413Z digest=sha256:9662de4ff9d7751d1cae62aeefe347e6bd68c00565ebab426b29be908ee8455e

Observation a5707012-71a3-481c-9bc5-2976e54a613d · outbound

This paper cites MMBench: Is Your Multi-modal Model an All-around Player?.

Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering MMBench: Is Your Multi-modal Model an All-around Player?

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-12T12:52:30.859366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:52:30.859366Z digest=sha256:aa76ee19460a726d396633dc3bfee02bfb7cf6977057bb77263ad57bebb5d8da

Observation a51a05f5-e4f0-4403-948f-bbb4bd55400d · outbound

This paper cites Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question Answering.

Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question Answering

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:52:31.637469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T12:52:30.864913Z digest=sha256:cba3fe784f1f7d371ea25c7b0cf1e4e230c84b495b8bdb2a33603e42ea0d27a0

Observation b04fdc93-75e9-4cf6-9c77-048753ca6730 · outbound

This paper cites Ok-VQA: A Visual Question Answering Benchmark Requiring External Knowledge.

Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering Ok-VQA: A Visual Question Answering Benchmark Requiring External Knowledge

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:52:31.623396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T12:52:30.869230Z digest=sha256:e0b5c39c4f8648b79d6a65f2ad24a93f7456785c6051312b9435561e29e29e8e

Observation 27ef76d9-b44b-46c6-9511-e5b546c06af8 · outbound

This paper cites MM1: Methods, Analy- sis & Insights from Multimodal LLM Pre-training.

Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering MM1: Methods, Analy- sis & Insights from Multimodal LLM Pre-training

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:52:31.608832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T12:52:30.873777Z digest=sha256:7f89732bf61f8d80d5cf26d00ccc3dc4a19d444fd92948897ab2866d91e7fc48

Observation 98a48acc-3c2c-46a0-8600-9995ff825a27 · outbound

This paper cites Encyclopedic VQA: Visual Questions About Detailed Properties of Fine-Grained Categories.

Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering Encyclopedic VQA: Visual Questions About Detailed Properties of Fine-Grained Categories

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:52:31.594478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T12:52:30.878115Z digest=sha256:71047fb187cf0bf6a591530d6d115bdb5613764937bdd272a411e39791a9fb1d

Observation 31935bd8-1fd6-4e7c-bb97-7bfc22cd9ec7 · outbound

This paper cites PlotQA: Reasoning over Scientific Plots.

Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering PlotQA: Reasoning over Scientific Plots

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:52:31.580820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T12:52:30.882608Z digest=sha256:327c4ba6c70b5c56c7cd2d1916a7a6f7502d040e6ee3880e4caa3c7cd3aba858

Observation 257e8f23-de4a-4de9-b5e5-68a594c6303a · outbound

This paper cites Revisiting Image Cap- tioning Training Paradigm via Direct CLIP-based Optimiza- tion.

Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering Revisiting Image Cap- tioning Training Paradigm via Direct CLIP-based Optimiza- tion

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:52:31.567477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T12:52:30.887116Z digest=sha256:0e117803891ca7b34f15045ed43cb88d0d1eee7d543307402093f297450a273e

Observation 155570d5-f38f-4430-840d-a3a0a51a6311 · outbound

This paper cites Training Language Models to Follow Instructions with Human Feedback.

Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering Training Language Models to Follow Instructions with Human Feedback

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:52:31.553654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T12:52:30.891795Z digest=sha256:478e58d5018c90fab4ebcdc5178a4a669db640339cd4602f88c814dbd1af0e02

Observation e5817235-ee81-48fb-9eb4-e46b09e94648 · outbound

This paper cites RoRA-VLM: Robust Retrieval-Augmented Vision Language Models.

Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering RoRA-VLM: Robust Retrieval-Augmented Vision Language Models

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-12T12:52:30.896241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:52:30.896241Z digest=sha256:2a58d455e7d05411e09f306681bc92bfee901ee8febc1ab4da504751f04591fb

Observation f4ca4721-3864-4d6d-bb24-2540961054c4 · outbound

This paper cites Learning Transferable Visual Models from Natural Language Supervi- sion.

Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering Learning Transferable Visual Models from Natural Language Supervi- sion

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:52:31.539693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T12:52:30.900945Z digest=sha256:b57b4c293e35ae6fdd9af613c6325ebb9866ba427b9251e4c3d1d09a09844a4f

Observation 3c97973b-06d1-4bfd-8cc8-61ce028e3368 · outbound

This paper cites Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer.

Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:52:31.524944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T12:52:30.905609Z digest=sha256:98c4e0462d05e2f7ba7f2b92c2e148d447d91ee437bf0a63f52afe88c5144dfe

Observation 703ed348-9f2d-4986-b079-dc40e27e3f5d · outbound

This paper cites In- Context Retrieval-Augmented Language Models.

Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering In- Context Retrieval-Augmented Language Models

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:52:31.511110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T12:52:30.910980Z digest=sha256:1e327def3b3b6bd0368c661f2b6a5f4f09c3b09d436dbeb775eec6e801b40909

Observation ce2f2d31-a14c-41ac-a703-1e7baf963715 · outbound

This paper cites A-OKVQA: A Benchmark for Visual Question Answering Using World Knowledge.

Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering A-OKVQA: A Benchmark for Visual Question Answering Using World Knowledge

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:52:31.496759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T12:52:30.915553Z digest=sha256:3a6b1b6a79cd7e1c024de16f3dd2d677ff3fa5bb437b4fe946fbb004f29113e0

Observation a7541de2-06fb-46d7-a623-70e4be65e50b · outbound

This paper cites KVQA: Knowledge-aware Visual Question Answering.

Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering KVQA: Knowledge-aware Visual Question Answering

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:52:31.482476Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T12:52:30.920299Z digest=sha256:64a0cba48b39a91573fa9f7169ce03e2a83ca4568c09f2b49450249ecdaf9e73

Observation d825cc38-08a9-49b3-be16-bfa5fa793f83 · outbound

This paper cites Towards VQA Models That Can Read.

Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering Towards VQA Models That Can Read

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:52:31.468289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T12:52:30.926189Z digest=sha256:541ec67024d624ce8077458b79079ff2ab711d75351d3c06748f4f055bc5c991

Observation b9ba3b35-fc43-4983-8e92-e5a8a5022844 · outbound

This paper cites WIT: Wikipedia-based Image Text Dataset for Multimodal Multilingual Machine Learning.

Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering WIT: Wikipedia-based Image Text Dataset for Multimodal Multilingual Machine Learning

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:52:31.453760Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T12:52:30.931685Z digest=sha256:89c88a0ba5f0011e41782a1b4377e90eb5b3c259687c03bd0c8b13237d14e027

Observation 604dbd24-601d-4acc-8380-8d19cf97cc55 · outbound

This paper cites EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters.

Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-12T12:52:30.936377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:52:30.936377Z digest=sha256:102001d71db697c7f820901d870cdeb5c44c9572528295e83b369829c868996d

Observation 446863de-a2e6-4ac4-8d64-66041767d9d4 · outbound

This paper cites Emu: Generative Pretraining in Multimodality.

Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering Emu: Generative Pretraining in Multimodality

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:52:31.438865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T12:52:30.941139Z digest=sha256:e5da00877bbc20afcbdde1e0bc3d74c8c80c05a3bc08733332b038760d8a2f58

Observation 3d914263-29fa-4526-bd81-7bbf64eb7de8 · outbound

This paper cites Stanford Alpaca: An Instruction-Following LLaMA Model, 2023.

Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering Stanford Alpaca: An Instruction-Following LLaMA Model, 2023

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:52:31.424330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T12:52:30.945605Z digest=sha256:c4b4d21edf0bee5770ac7989246ee8db7faf97ee945fbc46e8d7a4d7c33615c9

Observation b09f0594-c9af-487e-b06c-6714b6712f01 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering LLaMA: Open and Efficient Foundation Language Models

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-12T12:52:30.950194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:52:30.950194Z digest=sha256:7cc166341f34f4e5caf9307dd68ac6b56b8de200d1a96b59bac6a04759edb6b8

Observation ed0a2a25-31f6-433b-91d6-781c0281c5b8 · outbound

This paper cites InstructRetro: Instruction Tuning post Retrieval-Augmented Pretraining.

Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering InstructRetro: Instruction Tuning post Retrieval-Augmented Pretraining

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:52:31.407712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T12:52:30.954604Z digest=sha256:e8f25f00be3d4ab37ed2c4e04f326015ba6030bbf83cf180fce7f02b74c33b95

Observation 0e81e5b5-91de-415a-acc8-ba3e753f4aa6 · outbound

This paper cites UniIR: Training and Benchmarking Universal Multimodal Information Retrievers.

Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering UniIR: Training and Benchmarking Universal Multimodal Information Retrievers

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:52:31.392251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T12:52:30.959180Z digest=sha256:3bab318705c3cfd0ab4e5dfe3c139d67f67ea423ab3e81208dfb1e6530215564

Observation 493ee015-4639-4ca9-b37c-63b8946309e5 · outbound

This paper cites CLAIR-A: Leveraging Large Language Models to Judge Audio Captions.

Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering CLAIR-A: Leveraging Large Language Models to Judge Audio Captions

Reference 69

Resolution
verified exact
local_arxiv, observed 2026-08-12T12:52:31.064629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T12:52:30.963598Z digest=sha256:17a47c406e75d39593f9f9eb1676affaa51a75c096d38eab27ecede1b6a14f0d

Observation 56af160e-9fa7-4545-810e-5338275d0ded · outbound

This paper cites Grounding Language Models for Visual Entity Recognition.

Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering Grounding Language Models for Visual Entity Recognition

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:52:31.377396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T12:52:30.968344Z digest=sha256:42c41f7716fefac5fd863271fa8e5c4bbb8bd843e9a1aacb9d7fe91dcedc265f

Observation 39901545-cd60-41e3-ba81-fc266361067d · outbound

This paper cites EchoSight: Advancing Visual-Language Models with Wiki Knowledge.

Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering EchoSight: Advancing Visual-Language Models with Wiki Knowledge

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-12T12:52:30.972834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:52:30.972834Z digest=sha256:bf1c726e5e528bcd545fac091339ba2e95da08fe26c9a806f3963ba925ef9baa

Observation 2226b6c7-5663-403e-b31a-0275c673d69c · outbound

This paper cites mPLUG- Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration.

Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering mPLUG- Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:52:31.363083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T12:52:30.977585Z digest=sha256:ed493021b1cc2c4c77b4e5bfef4cb1685dd27ebb1ba291c65a805a4aaf41e7ef

Observation b59bc50d-3ed5-446c-9c78-48ade2622f41 · outbound

This paper cites RankRAG: Unifying Context Ranking with Retrieval-Augmented Generation in LLMs.

Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering RankRAG: Unifying Context Ranking with Retrieval-Augmented Generation in LLMs

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-12T12:52:30.982173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:52:30.982173Z digest=sha256:d8cb4f4c40e72f366459190715104b50520923e7f5d8857af593fbc364887558

Observation 255f2df9-65e1-405a-ba53-b125d817d9b6 · outbound

This paper cites Give a short answer.

Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering Give a short answer

Reference 74

Resolution
malformed identifier
raw_fallback, observed 2026-08-12T12:52:31.347818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T12:52:30.987353Z digest=sha256:cb36ff087fc5442d73bbb3a988667951bddf1fedc056a854f30e817e2c2e2e54

Observation 39fc22b7-cffb-4284-a342-eddaa3a21928 · outbound

This paper cites an unresolved cited work.

Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering Unresolved cited work

Reference 2024

Resolution
parse uncertain
raw_fallback, observed 2026-08-12T12:52:31.652208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T12:52:30.854808Z digest=sha256:8230f12fb230995554d805746df1fbecdd1c165ca58b5312d0c87e899df42096

Pith citing papers

Observation b7ea84fc-69d5-4fb1-baca-2570f8eedaaa · inbound

OMGM: Orchestrate Multiple Granularities and Modalities for Efficient Multimodal Retrieval cites this paper.

OMGM: Orchestrate Multiple Granularities and Modalities for Efficient Multimodal Retrieval Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T22:41:10.111823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T22:41:10.111823Z digest=sha256:e546a5156cb79d0253434e373811308049acbc46615ba165476a84e37cf264d2

Observation 8109059b-35dd-4671-8911-22da5803b3ed · inbound

Towards General Continuous Memory for Vision-Language Models cites this paper.

Towards General Continuous Memory for Vision-Language Models Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T14:46:44.356720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:46:44.356720Z digest=sha256:abeec9240d80cf76102ee0638cfd81cc3cb5e994c3f3f7828c1701450ee2d8c4

Observation da91c099-b84d-4551-8bd6-523da596ae4e · inbound

ExpStar: Towards Automatic Commentary Generation for Multi-discipline Scientific Experiments cites this paper.

ExpStar: Towards Automatic Commentary Generation for Multi-discipline Scientific Experiments Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T17:55:45.258885Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:55:45.258885Z digest=sha256:1607ce6a25bc83a870a076dcd4c1db1e22e588397db066de012d01574b6c6515

Observation 4882e389-ddcf-4108-8eca-840a641f684d · inbound

MathVis-Fine: Aligning Visual Supervision with Necessity via Progressive Dependency-Guided Training for Multimodal Mathematical Reasoning cites this paper.

MathVis-Fine: Aligning Visual Supervision with Necessity via Progressive Dependency-Guided Training for Multimodal Mathematical Reasoning Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-07-03T20:18:57.667225Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-06-27T01:23:40.564561Z digest=sha256:16a525aad79d3fa2b43152a5541fceed92d39718adc1a261b3417d598b4e5a7f