Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T23:53:05.141147Z
Paper Citation Record · LEDGER
As of 9 August 2026, this Paper Citation Record lists 23 of 23 outbound references and 0 inbound Pith citation observations for arXiv:2506.21596.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T23:53:05.141147Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
23 of 23 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation fae8a539-5ada-4a15-8b65-ee2a7d11eaa7 · outbound
Evaluating Multimodal Large Language Models on Educational Textbook Question Answering Visual instruction tuning,
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 00f0d1c9-a049-4ace-ba4c-bde3bd43e63d · outbound
Evaluating Multimodal Large Language Models on Educational Textbook Question Answering Improved Baselines with Visual Instruction Tuning
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4d946e54-d937-4ff2-b146-3688b3a99d89 · outbound
Evaluating Multimodal Large Language Models on Educational Textbook Question Answering The Llama 3 Herd of Models
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7bd79d0f-16e8-494e-a535-8c84e774f444 · outbound
Evaluating Multimodal Large Language Models on Educational Textbook Question Answering Taking the next step with generative artificial intelligence: The transformative role of multimodal large language models in science education,
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 7ce0486e-a7fe-4841-9a56-18d74747f6c9 · outbound
Evaluating Multimodal Large Language Models on Educational Textbook Question Answering On opportunities and challenges of large multimodal foundation models in education,
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 84789112-6207-439f-af96-0dcb1a5516bb · outbound
Evaluating Multimodal Large Language Models on Educational Textbook Question Answering Are you smarter than a sixth grader? textbook question an- swering for multimodal machine comprehension,
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 151a1e35-aeab-43f6-92c5-1e26822f2f62 · outbound
Evaluating Multimodal Large Language Models on Educational Textbook Question Answering A review on vision-language-based approaches: Challenges and applications,
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation c3a08743-84a5-47c9-9a33-7446f9f640aa · outbound
Evaluating Multimodal Large Language Models on Educational Textbook Question Answering Blip-2: Bootstrapping language- image pre-training with frozen image encoders and large language models,
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation f5eeea5c-cbeb-44a8-ae56-fd4e724b9905 · outbound
Evaluating Multimodal Large Language Models on Educational Textbook Question Answering Minigpt-4: Enhancing vision-language understanding with advanced large language models,
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 27591440-ac1b-49a3-ab2e-1d280f14405a · outbound
Evaluating Multimodal Large Language Models on Educational Textbook Question Answering Making the v in VQA matter: Elevating the role of image understanding in visual question answering,
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 8849f29d-7b38-406b-bb3d-46f2ce5e6da2 · outbound
Evaluating Multimodal Large Language Models on Educational Textbook Question Answering Ok-VQA: A visual question answering benchmark requiring external knowledge,
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation d620872c-86af-4ed9-9521-eeafac3d6cac · outbound
Evaluating Multimodal Large Language Models on Educational Textbook Question Answering Microsoft COCO: Common objects in context,
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation f6143b45-6167-4c8f-ac0a-421e7d3d17e2 · outbound
Evaluating Multimodal Large Language Models on Educational Textbook Question Answering SPIQA: A Dataset for Multimodal Question Answering on Scientific Papers
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0ea08724-8e05-4920-9e19-52d5ca82ee55 · outbound
Evaluating Multimodal Large Language Models on Educational Textbook Question Answering Eduvqa: A multimodal visual question answering framework for smart education,
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation d54e57bf-b35c-4e3e-a1f2-94a3233472c4 · outbound
Evaluating Multimodal Large Language Models on Educational Textbook Question Answering Enhancing textual textbook question answering with large language models and retrieval augmented generation,
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 57cf384d-eed3-4364-b65d-c3f0844c66a5 · outbound
Evaluating Multimodal Large Language Models on Educational Textbook Question Answering Enhancing textbook question answering with knowledge graph-augmented large language models,
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation f75ebdc2-61c3-4db1-8337-07a1e4745380 · outbound
Evaluating Multimodal Large Language Models on Educational Textbook Question Answering ISAAQ - mastering textbook questions with pre-trained transformers and bottom-up and top-down attention,
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 921db7cb-37bc-4d69-abda-1a6b2cb389f5 · outbound
Evaluating Multimodal Large Language Models on Educational Textbook Question Answering Imagebind: One embedding space to bind them all,
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 44b78eb2-0f23-4782-ab3a-a9a17102ef91 · outbound
Evaluating Multimodal Large Language Models on Educational Textbook Question Answering KDB.AI: The scalable vector database for ai,
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 0c0e506f-a6b0-4bfb-a07f-4847e7c6b754 · outbound
Evaluating Multimodal Large Language Models on Educational Textbook Question Answering GPT-4v(ision) system card,
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 8b00df41-b0d4-4458-a67e-8354fd5278a6 · outbound
Evaluating Multimodal Large Language Models on Educational Textbook Question Answering Gemini: A Family of Highly Capable Multimodal Models
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation af04fca8-030e-4bf4-95f4-b09c94cea7f5 · outbound
Evaluating Multimodal Large Language Models on Educational Textbook Question Answering Learning transferable visual models from natural language supervision,
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 8e02cfc0-9382-4039-afa1-b29ccdf59431 · outbound
Evaluating Multimodal Large Language Models on Educational Textbook Question Answering Vicuna: An open-source chatbot impressing gpt-4 with 90% chatgpt quality,
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
No inbound Pith citation observations are available.