Pith. sign in

Paper Citation Record · LEDGER

Evaluating Multimodal Large Language Models on Educational Textbook Question Answering

As of 19 August 2026, this Paper Citation Record lists 23 of 23 outbound references and 0 inbound Pith citation observations for arXiv:2506.21596.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.21596 v2

Coverage vector

measured 23 of 23 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T23:53:05.141147Z

measured 23 of 23 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

23 of 23 outbound references displayed

  • verified exact0
  • verified fuzzy18
  • unresolved5
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation fae8a539-5ada-4a15-8b65-ee2a7d11eaa7 · outbound

This paper cites Visual instruction tuning,.

Evaluating Multimodal Large Language Models on Educational Textbook Question Answering Visual instruction tuning,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T23:53:03.563420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:53:03.563420Z digest=sha256:2b9169c5ea96ac39893116c53383dc4b0eea24cd28b454a8a74568284664e462

Observation 00f0d1c9-a049-4ace-ba4c-bde3bd43e63d · outbound

This paper cites Improved Baselines with Visual Instruction Tuning.

Evaluating Multimodal Large Language Models on Educational Textbook Question Answering Improved Baselines with Visual Instruction Tuning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T23:53:03.623288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:53:03.623288Z digest=sha256:f527029c302abe1ef1b49929aa318ec2672e649479c837d80a078dd59a4cf59b

Observation 4d946e54-d937-4ff2-b146-3688b3a99d89 · outbound

This paper cites The Llama 3 Herd of Models.

Evaluating Multimodal Large Language Models on Educational Textbook Question Answering The Llama 3 Herd of Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T23:53:03.679170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:53:03.679170Z digest=sha256:ee56f58f944c6c8e91c09e6b4a7dcf9bd2e2a5ec5ad478d6c685b0e22fca9911

Observation 7bd79d0f-16e8-494e-a535-8c84e774f444 · outbound

This paper cites Taking the next step with generative artificial intelligence: The transformative role of multimodal large language models in science education,.

Evaluating Multimodal Large Language Models on Educational Textbook Question Answering Taking the next step with generative artificial intelligence: The transformative role of multimodal large language models in science education,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:53:09.040655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T23:53:03.747155Z digest=sha256:722fd7d2f2842d396cee60f8c4d3f0bc51707266b35ab27396efe47605749db1

Observation 7ce0486e-a7fe-4841-9a56-18d74747f6c9 · outbound

This paper cites On opportunities and challenges of large multimodal foundation models in education,.

Evaluating Multimodal Large Language Models on Educational Textbook Question Answering On opportunities and challenges of large multimodal foundation models in education,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:53:08.905573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T23:53:03.838513Z digest=sha256:36f634a968479cc573254b2ede0004815f055b443c1d782b1d4cd12871c75a1f

Observation 84789112-6207-439f-af96-0dcb1a5516bb · outbound

This paper cites Are you smarter than a sixth grader? textbook question an- swering for multimodal machine comprehension,.

Evaluating Multimodal Large Language Models on Educational Textbook Question Answering Are you smarter than a sixth grader? textbook question an- swering for multimodal machine comprehension,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:53:08.772329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T23:53:03.910260Z digest=sha256:828f0757d589a8a8d8bf2cb791de2f6fb67ac48af6254d8b26b1f557c9d87317

Observation 151a1e35-aeab-43f6-92c5-1e26822f2f62 · outbound

This paper cites A review on vision-language-based approaches: Challenges and applications,.

Evaluating Multimodal Large Language Models on Educational Textbook Question Answering A review on vision-language-based approaches: Challenges and applications,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:53:08.624559Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T23:53:03.979107Z digest=sha256:0aa955f849d6b427f8bcccbb02028397a5e41da40189df68f58572046450ad5d

Observation c3a08743-84a5-47c9-9a33-7446f9f640aa · outbound

This paper cites Blip-2: Bootstrapping language- image pre-training with frozen image encoders and large language models,.

Evaluating Multimodal Large Language Models on Educational Textbook Question Answering Blip-2: Bootstrapping language- image pre-training with frozen image encoders and large language models,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:53:08.480404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T23:53:04.026641Z digest=sha256:2b389a71f83d34cef336a840a7827375fa8698a1f9eca409a984eedbdf6fd9de

Observation f5eeea5c-cbeb-44a8-ae56-fd4e724b9905 · outbound

This paper cites Minigpt-4: Enhancing vision-language understanding with advanced large language models,.

Evaluating Multimodal Large Language Models on Educational Textbook Question Answering Minigpt-4: Enhancing vision-language understanding with advanced large language models,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:53:08.330726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T23:53:04.086129Z digest=sha256:b0d951d28aba9dcb55b20026ff68451d2fa6bdaf21d8199b29f137d19ed3322f

Observation 27591440-ac1b-49a3-ab2e-1d280f14405a · outbound

This paper cites Making the v in VQA matter: Elevating the role of image understanding in visual question answering,.

Evaluating Multimodal Large Language Models on Educational Textbook Question Answering Making the v in VQA matter: Elevating the role of image understanding in visual question answering,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:53:08.189944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T23:53:04.143669Z digest=sha256:f5ebe8c9faae274906d623fb23a9708d2bd4d55a9e02ba4f2165971c764a613f

Observation 8849f29d-7b38-406b-bb3d-46f2ce5e6da2 · outbound

This paper cites Ok-VQA: A visual question answering benchmark requiring external knowledge,.

Evaluating Multimodal Large Language Models on Educational Textbook Question Answering Ok-VQA: A visual question answering benchmark requiring external knowledge,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:53:08.013450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T23:53:04.184296Z digest=sha256:880bcc248743235cd2dafbbd5c24b56af6111f224042c2f0987aa4458095b82f

Observation d620872c-86af-4ed9-9521-eeafac3d6cac · outbound

This paper cites Microsoft COCO: Common objects in context,.

Evaluating Multimodal Large Language Models on Educational Textbook Question Answering Microsoft COCO: Common objects in context,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:53:07.691252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T23:53:04.276954Z digest=sha256:27b1c182c4422cc5d5e595946fcaddcfdaa93ee3cb890f1e803186da6fcc5241

Observation f6143b45-6167-4c8f-ac0a-421e7d3d17e2 · outbound

This paper cites SPIQA: A Dataset for Multimodal Question Answering on Scientific Papers.

Evaluating Multimodal Large Language Models on Educational Textbook Question Answering SPIQA: A Dataset for Multimodal Question Answering on Scientific Papers

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T23:53:04.357027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:53:04.357027Z digest=sha256:4b99fafb873a110f91aec6c88a2b0b5457eb9c9ba42b68f55efeae89f03dc316

Observation 0ea08724-8e05-4920-9e19-52d5ca82ee55 · outbound

This paper cites Eduvqa: A multimodal visual question answering framework for smart education,.

Evaluating Multimodal Large Language Models on Educational Textbook Question Answering Eduvqa: A multimodal visual question answering framework for smart education,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:53:07.417384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T23:53:04.425031Z digest=sha256:6fd34fdfc88f4c23cdc329c736c1231fe912cedbcc63a228284d2fbc14d3e40a

Observation d54e57bf-b35c-4e3e-a1f2-94a3233472c4 · outbound

This paper cites Enhancing textual textbook question answering with large language models and retrieval augmented generation,.

Evaluating Multimodal Large Language Models on Educational Textbook Question Answering Enhancing textual textbook question answering with large language models and retrieval augmented generation,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:53:07.151392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T23:53:04.508308Z digest=sha256:ec338abbb2d394ec5c4638e4246f9030236101b4f4c00bffbe9e7bb8b544e198

Observation 57cf384d-eed3-4364-b65d-c3f0844c66a5 · outbound

This paper cites Enhancing textbook question answering with knowledge graph-augmented large language models,.

Evaluating Multimodal Large Language Models on Educational Textbook Question Answering Enhancing textbook question answering with knowledge graph-augmented large language models,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:53:06.928385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T23:53:04.614008Z digest=sha256:988e57a5fe6c342ea9eb259326f5a7cc4e09d871b6b8b30cebc00ad35bf88fc7

Observation f75ebdc2-61c3-4db1-8337-07a1e4745380 · outbound

This paper cites ISAAQ - mastering textbook questions with pre-trained transformers and bottom-up and top-down attention,.

Evaluating Multimodal Large Language Models on Educational Textbook Question Answering ISAAQ - mastering textbook questions with pre-trained transformers and bottom-up and top-down attention,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:53:06.642236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T23:53:04.676803Z digest=sha256:ff3b819e713591f9514ea0f3471df721789022e764a6e48680ca8f46fe088ea4

Observation 921db7cb-37bc-4d69-abda-1a6b2cb389f5 · outbound

This paper cites Imagebind: One embedding space to bind them all,.

Evaluating Multimodal Large Language Models on Educational Textbook Question Answering Imagebind: One embedding space to bind them all,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:53:06.355230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T23:53:04.754929Z digest=sha256:f24d6505ec529ff45e008d7ceb0c2ce3165bab45da5ccff4af0785d21e63baa1

Observation 44b78eb2-0f23-4782-ab3a-a9a17102ef91 · outbound

This paper cites KDB.AI: The scalable vector database for ai,.

Evaluating Multimodal Large Language Models on Educational Textbook Question Answering KDB.AI: The scalable vector database for ai,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:53:06.116786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T23:53:04.846424Z digest=sha256:add76919aa66d86f0466302d6b9d04253323aa4e65d3b7c4dd75a406b94e6b51

Observation 0c0e506f-a6b0-4bfb-a07f-4847e7c6b754 · outbound

This paper cites GPT-4v(ision) system card,.

Evaluating Multimodal Large Language Models on Educational Textbook Question Answering GPT-4v(ision) system card,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:53:05.828269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T23:53:04.897266Z digest=sha256:6e23c40297d3e182174df898e4c86dde25af7d5c68a613a563abdfc570af8830

Observation 8b00df41-b0d4-4458-a67e-8354fd5278a6 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Evaluating Multimodal Large Language Models on Educational Textbook Question Answering Gemini: A Family of Highly Capable Multimodal Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T23:53:04.903476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:53:04.903476Z digest=sha256:ce98d20b3b327955f94c500120cca61c671c2f49ea9ac1700b012a43077700c5

Observation af04fca8-030e-4bf4-95f4-b09c94cea7f5 · outbound

This paper cites Learning transferable visual models from natural language supervision,.

Evaluating Multimodal Large Language Models on Educational Textbook Question Answering Learning transferable visual models from natural language supervision,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:53:05.582385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T23:53:04.997136Z digest=sha256:6ea97d7f0d85d6a7f44c65a66b9cf4d9a0034e892b733e459236cc6f5faee3a4

Observation 8e02cfc0-9382-4039-afa1-b29ccdf59431 · outbound

This paper cites Vicuna: An open-source chatbot impressing gpt-4 with 90% chatgpt quality,.

Evaluating Multimodal Large Language Models on Educational Textbook Question Answering Vicuna: An open-source chatbot impressing gpt-4 with 90% chatgpt quality,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:53:05.393712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T23:53:05.141147Z digest=sha256:fcb7a3e3991f6e9b15d395a509606240a4914998b4e4c35fc7e5412c6cc4c0fb

Pith citing papers

No inbound Pith citation observations are available.