Pith. sign in

Paper Citation Record · LEDGER

Enhancing Scientific Visual Question Answering through Multimodal Reasoning and Ensemble Modeling

As of 9 August 2026, this Paper Citation Record lists 26 of 26 outbound references and 1 inbound Pith citation observation for arXiv:2507.06183.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.06183 v1

Coverage vector

measured 26 of 26 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T19:13:26.851065Z

measured 27 of 27 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-10T18:31:15.525331Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T00:25:52.598353Z

Reference resolution

26 of 26 outbound references displayed

  • verified exact1
  • verified fuzzy0
  • unresolved24
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6da5d4b8-7592-4a38-9053-b9b525308c80 · outbound

This paper cites Qwen2.5-VL Technical Report.

Enhancing Scientific Visual Question Answering through Multimodal Reasoning and Ensemble Modeling Qwen2.5-VL Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T19:13:24.132881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:13:24.132881Z digest=sha256:264fb60bd3c85127a7567f8be8f9f9ac20ac4b1aaafc47077ef95127c0792845

Observation 5d27ca47-af39-4d2e-8e17-8919808a8380 · outbound

This paper cites an unresolved cited work.

Enhancing Scientific Visual Question Answering through Multimodal Reasoning and Ensemble Modeling Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-06T19:13:27.873009Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T19:13:24.213162Z digest=sha256:13391c410bb9366fa3f01a500d4310e6f96483f3b3efb9c25cbcd3bc0525085a

Observation 19b79cb6-d350-4e1a-b55f-207d2b621769 · outbound

This paper cites Chart-based Reasoning: Transferring Capabilities from LLMs to VLMs.

Enhancing Scientific Visual Question Answering through Multimodal Reasoning and Ensemble Modeling Chart-based Reasoning: Transferring Capabilities from LLMs to VLMs

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T19:13:24.333743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:13:24.333743Z digest=sha256:24ae7dbdee9de4f39c2674baa18cbfa7f60334ae28bd64b7db7edce448c241c4

Observation 651599da-b75b-4bf5-a8a6-0d9e8928bf6b · outbound

This paper cites ChartLlama: A Multimodal LLM for Chart Understanding and Generation.

Enhancing Scientific Visual Question Answering through Multimodal Reasoning and Ensemble Modeling ChartLlama: A Multimodal LLM for Chart Understanding and Generation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T19:13:24.426698Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:13:24.426698Z digest=sha256:3065b002f5ce0ecf9000dbae3c26689f05b5780cd1b6e9af42429d7817952f03

Observation e3cac431-4703-4012-a7d0-1aba6291043e · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

Enhancing Scientific Visual Question Answering through Multimodal Reasoning and Ensemble Modeling LoRA: Low-Rank Adaptation of Large Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T19:13:24.545409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:13:24.545409Z digest=sha256:c4de48f082d37887b84fcfa9bd116b154f9fad3b5beddd118c1ef88869b26d43

Observation 88c3bd8a-2ae0-4809-8960-838b84349693 · outbound

This paper cites Farhan Ishmam, Md.

Enhancing Scientific Visual Question Answering through Multimodal Reasoning and Ensemble Modeling Farhan Ishmam, Md

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T19:13:24.670393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:13:24.670393Z digest=sha256:3395b09e1d5dd77a665471359492a61f3ff2aece1ed59536444976d71c58cd96

Observation 38dfd1a9-4b6d-44c1-b93d-2e265a96ae96 · outbound

This paper cites A Comprehensive Survey on Visual Question Answering Datasets and Algorithms.

Enhancing Scientific Visual Question Answering through Multimodal Reasoning and Ensemble Modeling A Comprehensive Survey on Visual Question Answering Datasets and Algorithms

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T19:13:24.765236Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:13:24.765236Z digest=sha256:69724c10b8a6859cfb7098dcdea6208f29376285285df96969523b2f0afccca9

Observation eed876b4-30a5-41f5-aa5c-ebc3e83ca83d · outbound

This paper cites Efficient Memory Management for Large Language Model Serving with PagedAttention.

Enhancing Scientific Visual Question Answering through Multimodal Reasoning and Ensemble Modeling Efficient Memory Management for Large Language Model Serving with PagedAttention

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T19:13:24.886671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:13:24.886671Z digest=sha256:258707f9f598f0d9f476a5022d9296ffea5044e1b8f9ce0036d4ba3ab3e6aa99

Observation ff1b53f2-dcc8-4d63-bc0f-25b7544b3dcb · outbound

This paper cites MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts.

Enhancing Scientific Visual Question Answering through Multimodal Reasoning and Ensemble Modeling MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T19:13:24.964823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:13:24.964823Z digest=sha256:8994331d44e99ed43c13410c92b5d628aa6193368eb6bb263f4d775bb67bb63d

Observation b4289eba-bdfa-4cda-9eb2-9d3971879692 · outbound

This paper cites UniChart: A Universal Vision-language Pretrained Model for Chart Comprehension and Reasoning.

Enhancing Scientific Visual Question Answering through Multimodal Reasoning and Ensemble Modeling UniChart: A Universal Vision-language Pretrained Model for Chart Comprehension and Reasoning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T19:13:25.075905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:13:25.075905Z digest=sha256:0785b2dcdf4691bf0bbc0252c11c6941f160952ba07a9afd451b33a69b485977

Observation 6baaecf6-9a99-4e9e-94b2-f2c15aff8955 · outbound

This paper cites ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning.

Enhancing Scientific Visual Question Answering through Multimodal Reasoning and Ensemble Modeling ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T19:13:25.111434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:13:25.111434Z digest=sha256:bc4e26925de16bbbd8f3a0ff8038eab3cb18fb197428c7693d8450c61df89ca1

Observation d31e54cf-6ef2-420c-9a96-2ed9844b3045 · outbound

This paper cites Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs.

Enhancing Scientific Visual Question Answering through Multimodal Reasoning and Ensemble Modeling Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T19:13:25.235115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:13:25.235115Z digest=sha256:f7283d762d8fc7de8423f84ef59ae239e418c67f8b625193d608358970d31143

Observation 176e75ca-64bb-44ce-bbe3-d5ae04c98d91 · outbound

This paper cites SPIQA: A Dataset for Multimodal Question Answering on Scientific Papers.

Enhancing Scientific Visual Question Answering through Multimodal Reasoning and Ensemble Modeling SPIQA: A Dataset for Multimodal Question Answering on Scientific Papers

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T19:13:25.301379Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:13:25.301379Z digest=sha256:831c9616b007f8b4f420f7013c749fd27d3ffc12bf2c16e519a593d0772f6532

Observation 10d6402f-023e-44e9-8f6c-cbf7789d6985 · outbound

This paper cites Direct Preference Optimization: Your Language Model is Secretly a Reward Model.

Enhancing Scientific Visual Question Answering through Multimodal Reasoning and Ensemble Modeling Direct Preference Optimization: Your Language Model is Secretly a Reward Model

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T19:13:25.320183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:13:25.320183Z digest=sha256:a5f110f25f5c209dd490881f8c9ba0c6b440a3a1e197e834e4c9d98f3fe654a7

Observation a293879e-2f35-481d-8f63-f527c0635758 · outbound

This paper cites Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer.

Enhancing Scientific Visual Question Answering through Multimodal Reasoning and Ensemble Modeling Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T19:13:25.450751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:13:25.450751Z digest=sha256:f0d96d8876c5bbd0b6d2b50e73465bdf08efc4449d8cd4319c17129249ed7c2a

Observation de13ff66-f632-4326-be44-0280c4f4aa6b · outbound

This paper cites Exploring Rewriting Approaches for Different Conversational Tasks.

Enhancing Scientific Visual Question Answering through Multimodal Reasoning and Ensemble Modeling Exploring Rewriting Approaches for Different Conversational Tasks

Reference 16

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T19:13:27.260471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T19:13:25.559058Z digest=sha256:3eece2775a1374d350882b872f30c3e5da4cf33967dd5ee4fe1bb8fb1745569a

Observation 007578da-c4f2-4f92-89fe-3b8418d904da · outbound

This paper cites LlamaV-o1: Rethinking Step-by-step Visual Reasoning in LLMs.

Enhancing Scientific Visual Question Answering through Multimodal Reasoning and Ensemble Modeling LlamaV-o1: Rethinking Step-by-step Visual Reasoning in LLMs

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T19:13:25.672059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:13:25.672059Z digest=sha256:7017ee6bca3a5583e3a5aec1b4e5202d5915bee542bba059d2eefc10e86a0b08

Observation 7f3dd424-524f-45f1-b19c-3ed16265bb52 · outbound

This paper cites VL-Rethinker: Incentivizing Self-Reflection of Vision-Language Models with Reinforcement Learning.

Enhancing Scientific Visual Question Answering through Multimodal Reasoning and Ensemble Modeling VL-Rethinker: Incentivizing Self-Reflection of Vision-Language Models with Reinforcement Learning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T19:13:25.782023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:13:25.782023Z digest=sha256:667c4c39895ca8692bd00de17f23d6c48edd21c10e6e5f1df3e215948adc3673

Observation 2c086010-ccc6-4df4-bad9-24af0fb845d8 · outbound

This paper cites an unresolved cited work.

Enhancing Scientific Visual Question Answering through Multimodal Reasoning and Ensemble Modeling Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-06T19:13:27.670426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T19:13:25.915671Z digest=sha256:3db5c03c1c3612b4ab6c8e90d863dcc1b96355bb0aa6f1cc820968ae34c21a92

Observation c2e00931-b974-435b-9879-eb8708346140 · outbound

This paper cites Chain-of-Thought Prompting Elicits Reasoning in Large Language Models.

Enhancing Scientific Visual Question Answering through Multimodal Reasoning and Ensemble Modeling Chain-of-Thought Prompting Elicits Reasoning in Large Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T19:13:26.028106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:13:26.028106Z digest=sha256:50234debc080eee5f7c339905d645d9ef7a3ac3ccd7cd3d88dd9f925b5ef47ee

Observation a23e7867-472e-4edc-8f2a-5adf0d707c99 · outbound

This paper cites Tree of Thoughts: Deliberate Problem Solving with Large Language Models.

Enhancing Scientific Visual Question Answering through Multimodal Reasoning and Ensemble Modeling Tree of Thoughts: Deliberate Problem Solving with Large Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T19:13:26.141905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:13:26.141905Z digest=sha256:ac24ece99223821b441de63a484af538dc6a092ae5040b0a4c63177a7fdbdab0

Observation da7a07fe-dd6e-4d25-8be5-d100c8cd1546 · outbound

This paper cites SPRI: Aligning Large Language Models with Context-Situated Principles.

Enhancing Scientific Visual Question Answering through Multimodal Reasoning and Ensemble Modeling SPRI: Aligning Large Language Models with Context-Situated Principles

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-08-06T19:13:27.061026Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T19:13:26.256671Z digest=sha256:a2e5051a25a8e335e5d31984b15d608aea9f0d5d1afdabcd45e6a0505e8bd56c

Observation 9766eb6b-7fdc-418d-82cd-7f5121d6651e · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Enhancing Scientific Visual Question Answering through Multimodal Reasoning and Ensemble Modeling DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T19:13:26.374962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:13:26.374962Z digest=sha256:c300379ed161ad60093001f536b65d5dc55167eb15c48967995612ab92d243af

Observation f79664c0-fd49-4a2b-9e86-14d2de44c813 · outbound

This paper cites InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models.

Enhancing Scientific Visual Question Answering through Multimodal Reasoning and Ensemble Modeling InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T19:13:26.528077Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:13:26.528077Z digest=sha256:cc485d81aeeb29e6dee4618899d945362928b6981808f21198b19744fdc786dd

Observation 4232198f-c475-4031-a872-c21921279274 · outbound

This paper cites URL: " 'urlintro :=.

Enhancing Scientific Visual Question Answering through Multimodal Reasoning and Ensemble Modeling URL: " 'urlintro :=

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T19:13:26.732433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:13:26.732433Z digest=sha256:efe69e739af1d3ca7b81da90815f02aec6a57d9ad708d3afd665a77ae1382af1

Observation a0195551-95c0-4e78-b3de-2f939d32a89e · outbound

This paper cites write newline.

Enhancing Scientific Visual Question Answering through Multimodal Reasoning and Ensemble Modeling write newline

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T19:13:26.851065Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:13:26.851065Z digest=sha256:7d1d7843a4869af6c231c1e2f4b055bd7cc60f6636d801d61826ea7f478c359c

Pith citing papers

Observation c81f5110-650b-476d-88aa-b2734b237dd2 · inbound

EpiBench: Benchmarking Multi-turn Research Workflows for Multimodal Agents cites this paper.

EpiBench: Benchmarking Multi-turn Research Workflows for Multimodal Agents Enhancing Scientific Visual Question Answering through Multimodal Reasoning and Ensemble Modeling

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:25:52.617911Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T18:31:15.525331Z digest=sha256:a2acc8216ce76ca3005abea4aa38a85792f243a94b23b480871877e27b4eff22