Pith. sign in

Paper Citation Record · LEDGER

MagiC: Evaluating Multimodal Cognition Toward Grounded Visual Reasoning

As of 18 August 2026, this Paper Citation Record lists 39 of 39 outbound references and 0 inbound Pith citation observations for arXiv:2507.07297.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.07297 v1

Coverage vector

measured 39 of 39 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T18:51:06.202263Z

measured 39 of 39 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

39 of 39 outbound references displayed

  • verified exact4
  • verified fuzzy12
  • unresolved23
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 01079c24-8587-41b4-9ee6-5a8b73321352 · outbound

This paper cites Llama 3 model card.

MagiC: Evaluating Multimodal Cognition Toward Grounded Visual Reasoning Llama 3 model card

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:03.420164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:51:03.420164Z digest=sha256:a4a2a32821356a1f84760df04b0c168e028e589bd73d8d945392f6081a2383f8

Observation 4adc7ce3-70be-4142-8439-c3db538f8671 · outbound

This paper cites Introducing the next generation of claude.

MagiC: Evaluating Multimodal Cognition Toward Grounded Visual Reasoning Introducing the next generation of claude

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:51:08.910433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-06T18:51:03.478021Z digest=sha256:f92e7390e34d04568ae4a844769a1b2d6b10f80a073f96c1a06abe07b1271f73

Observation a287fb6c-90de-4f06-9964-f737307741be · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

MagiC: Evaluating Multimodal Cognition Toward Grounded Visual Reasoning Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:03.544169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:51:03.544169Z digest=sha256:cd3fbe12e52c08b3674fb8e23ed7a38e7ab534e15ccc5e4e1a25a2b17fdb918d

Observation e5658769-623e-418d-8a07-f020d37f21a0 · outbound

This paper cites Qwen2.5-VL Technical Report.

MagiC: Evaluating Multimodal Cognition Toward Grounded Visual Reasoning Qwen2.5-VL Technical Report

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:03.629099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:51:03.629099Z digest=sha256:aaaf16b54d3678acbeb730a8e25b081596f49c9863e12e6eeac649d7c06f203e

Observation 847a7c53-0ffc-4be9-b655-69daa69edf73 · outbound

This paper cites AiR : Attention with reasoning capability.

MagiC: Evaluating Multimodal Cognition Toward Grounded Visual Reasoning AiR : Attention with reasoning capability

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:51:08.713385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-06T18:51:03.686648Z digest=sha256:175a874aca5b7b41b46167605b97387b7c5305c5c428fb6d281d21ab04e0c2ea

Observation 26524540-6843-4f25-8d64-d6d90b3161f9 · outbound

This paper cites Measuring and improving chain-of-thought reasoning in vision-language models.

MagiC: Evaluating Multimodal Cognition Toward Grounded Visual Reasoning Measuring and improving chain-of-thought reasoning in vision-language models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:03.766806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:51:03.766806Z digest=sha256:820d0a3c5cc1722fc9f3d0223331764d6033538d7b98f4129a88ab087d8f45f3

Observation f50ac4c8-1124-4a79-b6c8-e32f1c9947e9 · outbound

This paper cites See, Think, Confirm: Interactive Prompting Between Vision and Language Models for Knowledge-based Visual Reasoning.

MagiC: Evaluating Multimodal Cognition Toward Grounded Visual Reasoning See, Think, Confirm: Interactive Prompting Between Vision and Language Models for Knowledge-based Visual Reasoning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:03.841863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:51:03.841863Z digest=sha256:6734ce7eac25d64894a6bae5f78d131c01f9876cd6679c899e8aa4c2c5adce4c

Observation f08cb58b-885c-473a-b805-4942db340f1d · outbound

This paper cites Spatialrgpt: Grounded spatial reasoning in vision-language models.

MagiC: Evaluating Multimodal Cognition Toward Grounded Visual Reasoning Spatialrgpt: Grounded spatial reasoning in vision-language models

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:51:08.565046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-06T18:51:03.907226Z digest=sha256:22ae8118e03ea06e26dea01eed86e28721717555e3f3caec38f2e69364654434

Observation 4de4e951-11d1-4b5b-8837-207ab0b1cd12 · outbound

This paper cites Aya-vision model card.

MagiC: Evaluating Multimodal Cognition Toward Grounded Visual Reasoning Aya-vision model card

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:51:08.411182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-06T18:51:03.976013Z digest=sha256:509046f52c1e08fe93ffd2d607addbf0ef6154c3e75d9c696b5b740100493312

Observation 9428915f-571b-493d-8bd4-4e9c9f3684c4 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

MagiC: Evaluating Multimodal Cognition Toward Grounded Visual Reasoning Gemini: A Family of Highly Capable Multimodal Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:04.037683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:51:04.037683Z digest=sha256:996785b673d705285bd1a805dd6739dfb8978638a82c4d663163adace8f3a977

Observation 803f8c99-c8b1-4ea9-aa9e-dd8cd31a484e · outbound

This paper cites Improved Visual Grounding through Self-Consistent Explanations.

MagiC: Evaluating Multimodal Cognition Toward Grounded Visual Reasoning Improved Visual Grounding through Self-Consistent Explanations

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-08-06T18:51:07.033159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-06T18:51:04.109878Z digest=sha256:0c1efaeaf1f313c4adcffeddc4c3192ca8d77c3dbdf8acf77b4130bed87cbbc0

Observation a94447db-402c-4f07-9daf-f7be8ed79e9b · outbound

This paper cites Gqa: A new dataset for real-world visual reasoning and compositional question answering.

MagiC: Evaluating Multimodal Cognition Toward Grounded Visual Reasoning Gqa: A new dataset for real-world visual reasoning and compositional question answering

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:51:08.229609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-06T18:51:04.201246Z digest=sha256:2f819af15a2d33c4b63483962caf0b3c1b9fc13b44d46e0b5bf026ef0c448949

Observation 4bfa830a-d607-4b6d-a050-a712d764dd7a · outbound

This paper cites GPT-4o System Card.

MagiC: Evaluating Multimodal Cognition Toward Grounded Visual Reasoning GPT-4o System Card

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:04.289982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:51:04.289982Z digest=sha256:0d1269207c976556f41ba9134fc647b3d2ebfce3124582d87f8a4cdfdd28ddc6

Observation c891e4df-803f-47e8-89ec-cdb9b2b70354 · outbound

This paper cites Weakly supervised grounding for vqa in vision-language transformers.

MagiC: Evaluating Multimodal Cognition Toward Grounded Visual Reasoning Weakly supervised grounding for vqa in vision-language transformers

Reference 14

Resolution
verified exact
doi, observed 2026-08-06T18:51:06.522495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-06T18:51:04.365691Z digest=sha256:49bffddd867194a7fdf7e8dc7a6fa383ae575094b6b0416b26aca4c9ff12beba

Observation 89a1cba5-0ab5-4c6a-8fa0-fe426fcdecb7 · outbound

This paper cites SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension.

MagiC: Evaluating Multimodal Cognition Toward Grounded Visual Reasoning SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:04.420426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:51:04.420426Z digest=sha256:ed2633b88a20457cb71ac035ee4bf29f9349d1a7a8d40fc8f23ffa52094f35fa

Observation 2cc45b3f-a014-47a5-9625-f7b95e3a6b91 · outbound

This paper cites BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models.

MagiC: Evaluating Multimodal Cognition Toward Grounded Visual Reasoning BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:04.468078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:51:04.468078Z digest=sha256:45ee3a856598236a6aca6e83ca6e812fc71134f5bcda00c64dd999c13692168d

Observation bc1c45f8-a047-46a9-a055-0e97c45ce61a · outbound

This paper cites Visual instruction tuning.

MagiC: Evaluating Multimodal Cognition Toward Grounded Visual Reasoning Visual instruction tuning

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:51:08.068541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-06T18:51:04.531795Z digest=sha256:694ee10c6004a819f090979ef556c798fc6d3f1d57364d00f48cb423b2108ea1

Observation 6de4a6e7-42da-4ea2-888b-26422013531a · outbound

This paper cites an unresolved cited work.

MagiC: Evaluating Multimodal Cognition Toward Grounded Visual Reasoning Unresolved cited work

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:04.586706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:51:04.586706Z digest=sha256:92357fc7d6ed99339fb921e35e0a530d6ef8d52eb66a3b430c477e6060553b8f

Observation 7d3d5810-42b4-4185-abcb-0427d493b7b5 · outbound

This paper cites Deepseek-vl: Towards real-world vision-language understanding, 2024.

MagiC: Evaluating Multimodal Cognition Toward Grounded Visual Reasoning Deepseek-vl: Towards real-world vision-language understanding, 2024

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:04.654427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:51:04.654427Z digest=sha256:0e98e475aeb6e8b3fe05cca3c627e4b4eca360fae223eb0631751464c5635938

Observation 3326c16b-02c8-4888-a0f1-2431d76235b4 · outbound

This paper cites Whiteboard-of-Thought: Thinking Step-by-Step Across Modalities.

MagiC: Evaluating Multimodal Cognition Toward Grounded Visual Reasoning Whiteboard-of-Thought: Thinking Step-by-Step Across Modalities

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:04.723812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:51:04.723812Z digest=sha256:7f82bd39bceef12a9e7a459d9a96b65da44112608b9637243bfa479e5a1df776

Observation de87c41c-c010-4906-bd35-c2f9d1377340 · outbound

This paper cites Gpt-4v(ision) system card.

MagiC: Evaluating Multimodal Cognition Toward Grounded Visual Reasoning Gpt-4v(ision) system card

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:04.807420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:51:04.807420Z digest=sha256:4301ce42e6632350605751f6148571abc23dc2ae2a4cc5ec7528f7d6faf51b9e

Observation ced51df7-21a3-40f8-bc0b-9a2d7e1cd7f0 · outbound

This paper cites Introducing gpt-4.1 in the api.

MagiC: Evaluating Multimodal Cognition Toward Grounded Visual Reasoning Introducing gpt-4.1 in the api

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:51:07.933231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-06T18:51:04.872081Z digest=sha256:692ffbc398c97f23903e887bdc43b4be6acb9df385316772fdd997b4dc540b66

Observation 73f51922-0f6a-4e14-b238-55e7ff329eb9 · outbound

This paper cites Qvq: To see the world with wisdom, December 2024.

MagiC: Evaluating Multimodal Cognition Toward Grounded Visual Reasoning Qvq: To see the world with wisdom, December 2024

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:51:07.766881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-06T18:51:04.903213Z digest=sha256:ecc0137c85dce3c774676776a4d672e97b286fa97a547c8661f8741460dfdc5e

Observation 52598951-ae8e-4590-9c5c-23e5980ed9d9 · outbound

This paper cites Uncovering the full potential of visual grounding methods in VQA.

MagiC: Evaluating Multimodal Cognition Toward Grounded Visual Reasoning Uncovering the full potential of visual grounding methods in VQA

Reference 24

Resolution
verified exact
doi, observed 2026-08-06T18:51:06.383863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-06T18:51:04.962971Z digest=sha256:7ece04435bcd588cb673db35d053d4472e7cd3fabbcb10536a3f0bdd84cb7062

Observation e1e41564-69ac-402c-b641-dd779435f174 · outbound

This paper cites Visual Chain of Thought: Bridging Logical Gaps with Multimodal Infillings.

MagiC: Evaluating Multimodal Cognition Toward Grounded Visual Reasoning Visual Chain of Thought: Bridging Logical Gaps with Multimodal Infillings

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:05.057773Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:51:05.057773Z digest=sha256:a34d1828b08bab82fa2bd802b1e5f72ae84a7746875d892fa2e32d60465d967c

Observation 7264deb3-fdc8-413f-8112-00ef15330c3f · outbound

This paper cites Visual CoT: Advancing Multi-Modal Language Models with a Comprehensive Dataset and Benchmark for Chain-of-Thought Reasoning.

MagiC: Evaluating Multimodal Cognition Toward Grounded Visual Reasoning Visual CoT: Advancing Multi-Modal Language Models with a Comprehensive Dataset and Benchmark for Chain-of-Thought Reasoning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:05.113854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:51:05.113854Z digest=sha256:a43bb761b4e16f97f67c4c2c95dd26e228ee533d1839e1727ef8f0333a5130a2

Observation 3e314ea3-e67d-44e2-9326-95c0ef7349ea · outbound

This paper cites Gemma 3 Technical Report.

MagiC: Evaluating Multimodal Cognition Toward Grounded Visual Reasoning Gemma 3 Technical Report

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:05.213221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:51:05.213221Z digest=sha256:724c86d5b9b8e0278766cdb8238689d8881ceb2b0bf670eb31dd9b56fc948e4b

Observation 391d53de-c277-45ab-848a-cf4e971f8c19 · outbound

This paper cites LlamaV-o1: Rethinking Step-by-step Visual Reasoning in LLMs.

MagiC: Evaluating Multimodal Cognition Toward Grounded Visual Reasoning LlamaV-o1: Rethinking Step-by-step Visual Reasoning in LLMs

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:05.257142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:51:05.257142Z digest=sha256:58ed8fdacbb7fd936f25b27719058dfb3b33979c1b85d0d3f47daaec6dd35592

Observation 51ec8827-b6fb-4c75-a226-2429fe40ba5e · outbound

This paper cites Contrastive region guidance: Improving grounding in vision-language models without training.

MagiC: Evaluating Multimodal Cognition Toward Grounded Visual Reasoning Contrastive region guidance: Improving grounding in vision-language models without training

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:05.326894Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:51:05.326894Z digest=sha256:44d5f1bd6b0e9ba4ef5de916d82c19b56c634787d7e4d37c8636642434681487

Observation 2fa00a24-98fb-41ab-a6ca-c0a1990eabe7 · outbound

This paper cites Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization.

MagiC: Evaluating Multimodal Cognition Toward Grounded Visual Reasoning Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:05.402887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:51:05.402887Z digest=sha256:59d0573b2f724c084bcc1109d2d242adaf6bb599869791b3df86f5a8675503dc

Observation 2430184c-629d-4eec-8bc3-f9736569496e · outbound

This paper cites VISCO: Benchmarking Fine-Grained Critique and Correction Towards Self-Improvement in Visual Reasoning.

MagiC: Evaluating Multimodal Cognition Toward Grounded Visual Reasoning VISCO: Benchmarking Fine-Grained Critique and Correction Towards Self-Improvement in Visual Reasoning

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-08-06T18:51:06.738821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-06T18:51:05.455071Z digest=sha256:fdacace86b1312fa23c4315dccd8276e16f337d18cc1c724dde81e33f88ec421

Observation 62e285fa-da03-4007-b4f9-e48be3f54dcf · outbound

This paper cites Llava-onevision-chat: Improving chat with preference learning, September 2024.

MagiC: Evaluating Multimodal Cognition Toward Grounded Visual Reasoning Llava-onevision-chat: Improving chat with preference learning, September 2024

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:51:07.629867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-06T18:51:05.525846Z digest=sha256:5ec95f190d6d9f84e2c03fdb3da8cf79a59387db9520d24d8d822a9fc9da19f1

Observation b0d75234-27f1-4d5f-b1e7-762272879e96 · outbound

This paper cites LLaVA-CoT: Let Vision Language Models Reason Step-by-Step.

MagiC: Evaluating Multimodal Cognition Toward Grounded Visual Reasoning LLaVA-CoT: Let Vision Language Models Reason Step-by-Step

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:05.654601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:51:05.654601Z digest=sha256:3fca4d9736683a449a633ce608311e70298da68a1eda044c1b646851d386ead5

Observation 08cec980-e5df-487c-a7a2-2e261c508c7b · outbound

This paper cites Improving Visual Grounding by Encouraging Consistent Gradient-based Explanations.

MagiC: Evaluating Multimodal Cognition Toward Grounded Visual Reasoning Improving Visual Grounding by Encouraging Consistent Gradient-based Explanations

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:05.755348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:51:05.755348Z digest=sha256:9b66f3473922ef2f9ae3d698a1a02d707541eacff106ceb2a68d922a458d9d99

Observation a99c0d05-b26f-4792-a42a-7dc8685f882a · outbound

This paper cites Mm-vet: evaluating large multimodal models for integrated capabilities.

MagiC: Evaluating Multimodal Cognition Toward Grounded Visual Reasoning Mm-vet: evaluating large multimodal models for integrated capabilities

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:51:07.478808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-06T18:51:05.824802Z digest=sha256:328fb7d7f0854800d4cc3a0d7a10e9eaf6fd77da121d864f950b76b334d057c6

Observation 3d4c7102-8b8d-4b1d-b41b-10db977a64f9 · outbound

This paper cites Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi.

MagiC: Evaluating Multimodal Cognition Toward Grounded Visual Reasoning Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:51:07.321052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-06T18:51:05.946191Z digest=sha256:a599a05409e8db27243978b563e8c8c883c672cf17138877b90eb9a9c462bef1

Observation ca938052-b05d-4806-b200-a1166d0a1387 · outbound

This paper cites Improve Vision Language Model Chain-of-thought Reasoning.

MagiC: Evaluating Multimodal Cognition Toward Grounded Visual Reasoning Improve Vision Language Model Chain-of-thought Reasoning

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:06.064596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:51:06.064596Z digest=sha256:90f24634bd9df419e064d878fd522fabcd4de6ea92497afdd0e1d6e176b9790a

Observation 5cc0ad3c-9096-45f9-bb46-c620dd382045 · outbound

This paper cites Multimodal chain-of-thought reasoning in language models.

MagiC: Evaluating Multimodal Cognition Toward Grounded Visual Reasoning Multimodal chain-of-thought reasoning in language models

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:51:07.198965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-06T18:51:06.133863Z digest=sha256:558de7df6ac9f03b0364e0b616ed4ba373acb5a0381eef686931647efd70743e

Observation 6a330a3f-c247-42d8-9270-a560a91fd405 · outbound

This paper cites CoT-VLA: Visual Chain-of-Thought Reasoning for Vision-Language-Action Models.

MagiC: Evaluating Multimodal Cognition Toward Grounded Visual Reasoning CoT-VLA: Visual Chain-of-Thought Reasoning for Vision-Language-Action Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:06.202263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:51:06.202263Z digest=sha256:47c47e1182926f19438dd756c6552c03e5cf9b55e168c86e6e5b9575f6e4f7c8

Pith citing papers

No inbound Pith citation observations are available.