Pith. sign in

Paper Citation Record · LEDGER

MagiC: Evaluating Multimodal Cognition Toward Grounded Visual Reasoning

As of 7 August 2026, this Paper Citation Record lists 39 of 39 outbound references and 0 inbound Pith citation observations for arXiv:2507.07297.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.07297 v1

Coverage vector

measured 39 of 39 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T18:51:06.202263Z

measured 39 of 39 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

39 of 39 outbound references displayed

  • verified exact4
  • verified fuzzy12
  • unresolved23
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 01079c24-8587-41b4-9ee6-5a8b73321352 · outbound

This paper cites Llama 3 model card.

MagiC: Evaluating Multimodal Cognition Toward Grounded Visual Reasoning Llama 3 model card

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:03.420164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:51:03.420164Z digest=sha256:149ac05b85ff4daf5c531a8f693583a000bf082d2cbb6bbfbc56271a55207adb

Observation 4adc7ce3-70be-4142-8439-c3db538f8671 · outbound

This paper cites Introducing the next generation of claude.

MagiC: Evaluating Multimodal Cognition Toward Grounded Visual Reasoning Introducing the next generation of claude

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:51:08.910433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T18:51:03.478021Z digest=sha256:98be3954e30b61fa773133a7af5a02ef4aeabe18c1f89b3c3a3135b39d8e0ebb

Observation a287fb6c-90de-4f06-9964-f737307741be · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

MagiC: Evaluating Multimodal Cognition Toward Grounded Visual Reasoning Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:03.544169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:51:03.544169Z digest=sha256:3b678e8789b31d0e0b2192a61d2d693ac7486a3748e9158e1b303e028421863e

Observation e5658769-623e-418d-8a07-f020d37f21a0 · outbound

This paper cites Qwen2.5-VL Technical Report.

MagiC: Evaluating Multimodal Cognition Toward Grounded Visual Reasoning Qwen2.5-VL Technical Report

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:03.629099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:51:03.629099Z digest=sha256:ae552fcb8550f6241d863f985b2218562bdcd640133e99ce16625d26a71c6889

Observation 847a7c53-0ffc-4be9-b655-69daa69edf73 · outbound

This paper cites AiR : Attention with reasoning capability.

MagiC: Evaluating Multimodal Cognition Toward Grounded Visual Reasoning AiR : Attention with reasoning capability

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:51:08.713385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T18:51:03.686648Z digest=sha256:3cf50db903308e2050cd0ad5d0ab9b0e16eedebf202716c14d221c8210bae020

Observation 26524540-6843-4f25-8d64-d6d90b3161f9 · outbound

This paper cites Measuring and improving chain-of-thought reasoning in vision-language models.

MagiC: Evaluating Multimodal Cognition Toward Grounded Visual Reasoning Measuring and improving chain-of-thought reasoning in vision-language models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:03.766806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:51:03.766806Z digest=sha256:fc6daa90ae35f0c3b70092c17bba845a332febfbcfec7ff7399815d43d35b3f7

Observation f50ac4c8-1124-4a79-b6c8-e32f1c9947e9 · outbound

This paper cites See, Think, Confirm: Interactive Prompting Between Vision and Language Models for Knowledge-based Visual Reasoning.

MagiC: Evaluating Multimodal Cognition Toward Grounded Visual Reasoning See, Think, Confirm: Interactive Prompting Between Vision and Language Models for Knowledge-based Visual Reasoning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:03.841863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:51:03.841863Z digest=sha256:4596fe8a9e23496b59614cbb7f94422e39d321ce275ba67d3addf8585c3462a2

Observation f08cb58b-885c-473a-b805-4942db340f1d · outbound

This paper cites Spatialrgpt: Grounded spatial reasoning in vision-language models.

MagiC: Evaluating Multimodal Cognition Toward Grounded Visual Reasoning Spatialrgpt: Grounded spatial reasoning in vision-language models

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:51:08.565046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T18:51:03.907226Z digest=sha256:eb0a1383614ac500df1f7bde52462f2a9f9e753fe9814209186089c9a7618e4d

Observation 4de4e951-11d1-4b5b-8837-207ab0b1cd12 · outbound

This paper cites Aya-vision model card.

MagiC: Evaluating Multimodal Cognition Toward Grounded Visual Reasoning Aya-vision model card

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:51:08.411182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T18:51:03.976013Z digest=sha256:b8deb9c759d65961ae8e9fe7f6e96fd4d827c511cf235806b63557ab1690de6e

Observation 9428915f-571b-493d-8bd4-4e9c9f3684c4 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

MagiC: Evaluating Multimodal Cognition Toward Grounded Visual Reasoning Gemini: A Family of Highly Capable Multimodal Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:04.037683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:51:04.037683Z digest=sha256:bb4c8c80f235abad1c173248b396a8eda11631628805bffa5abf7996c0f55b0a

Observation 803f8c99-c8b1-4ea9-aa9e-dd8cd31a484e · outbound

This paper cites Improved Visual Grounding through Self-Consistent Explanations.

MagiC: Evaluating Multimodal Cognition Toward Grounded Visual Reasoning Improved Visual Grounding through Self-Consistent Explanations

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-08-06T18:51:07.033159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T18:51:04.109878Z digest=sha256:ec9317c7576660a2934e46e8bb8917c2f2e9a73731bbd2f5d2d10e62d8285475

Observation a94447db-402c-4f07-9daf-f7be8ed79e9b · outbound

This paper cites Gqa: A new dataset for real-world visual reasoning and compositional question answering.

MagiC: Evaluating Multimodal Cognition Toward Grounded Visual Reasoning Gqa: A new dataset for real-world visual reasoning and compositional question answering

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:51:08.229609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T18:51:04.201246Z digest=sha256:9e84194df70944091d8a31b0a06f63afd4ef156eb724278bb39c0e70ca375917

Observation 4bfa830a-d607-4b6d-a050-a712d764dd7a · outbound

This paper cites GPT-4o System Card.

MagiC: Evaluating Multimodal Cognition Toward Grounded Visual Reasoning GPT-4o System Card

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:04.289982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:51:04.289982Z digest=sha256:cc2ac86bb260f3f8dcd1e4fc31f36c5e6c4f759342523abf7f9edb30e6958d99

Observation c891e4df-803f-47e8-89ec-cdb9b2b70354 · outbound

This paper cites Weakly supervised grounding for vqa in vision-language transformers.

MagiC: Evaluating Multimodal Cognition Toward Grounded Visual Reasoning Weakly supervised grounding for vqa in vision-language transformers

Reference 14

Resolution
verified exact
doi, observed 2026-08-06T18:51:06.522495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T18:51:04.365691Z digest=sha256:6721ab4ad38c95e320c3bbb17b36f255e3ce581fe3d8d057c3fa05fc56e32192

Observation 89a1cba5-0ab5-4c6a-8fa0-fe426fcdecb7 · outbound

This paper cites SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension.

MagiC: Evaluating Multimodal Cognition Toward Grounded Visual Reasoning SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:04.420426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:51:04.420426Z digest=sha256:d949eac0acd4075599871fab911767b4124682dabcd2336d0a3f2c492554797f

Observation 2cc45b3f-a014-47a5-9625-f7b95e3a6b91 · outbound

This paper cites BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models.

MagiC: Evaluating Multimodal Cognition Toward Grounded Visual Reasoning BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:04.468078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:51:04.468078Z digest=sha256:0f66567142bef31aadeaa1138cd37b6013ab6e6087a7a6cd705b2b32a214f01c

Observation bc1c45f8-a047-46a9-a055-0e97c45ce61a · outbound

This paper cites Visual instruction tuning.

MagiC: Evaluating Multimodal Cognition Toward Grounded Visual Reasoning Visual instruction tuning

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:51:08.068541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T18:51:04.531795Z digest=sha256:cdf5e1201aee342185c1b43ae00e445275572f1238cb0261c3ae61a9335c67a8

Observation 6de4a6e7-42da-4ea2-888b-26422013531a · outbound

This paper cites an unresolved cited work.

MagiC: Evaluating Multimodal Cognition Toward Grounded Visual Reasoning Unresolved cited work

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:04.586706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:51:04.586706Z digest=sha256:885fc82a9d82cb44c14466dd2bf9508053741bc4f338875c24c23b03862dbd80

Observation 7d3d5810-42b4-4185-abcb-0427d493b7b5 · outbound

This paper cites Deepseek-vl: Towards real-world vision-language understanding, 2024.

MagiC: Evaluating Multimodal Cognition Toward Grounded Visual Reasoning Deepseek-vl: Towards real-world vision-language understanding, 2024

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:04.654427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:51:04.654427Z digest=sha256:88afb6b6a0de3934ccb3fc821887f4579f1b159eb424f7baaa5e0000ca26cdc6

Observation 3326c16b-02c8-4888-a0f1-2431d76235b4 · outbound

This paper cites Whiteboard-of-Thought: Thinking Step-by-Step Across Modalities.

MagiC: Evaluating Multimodal Cognition Toward Grounded Visual Reasoning Whiteboard-of-Thought: Thinking Step-by-Step Across Modalities

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:04.723812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:51:04.723812Z digest=sha256:02eab6620be182b322e387c7e2a0589c19b559da4a211c5837a818acbce8ce7b

Observation de87c41c-c010-4906-bd35-c2f9d1377340 · outbound

This paper cites Gpt-4v(ision) system card.

MagiC: Evaluating Multimodal Cognition Toward Grounded Visual Reasoning Gpt-4v(ision) system card

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:04.807420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:51:04.807420Z digest=sha256:cb5014e031c4e8d6206b04b43030c28c889ccc1e1db3b64f78e03bbe8034c603

Observation ced51df7-21a3-40f8-bc0b-9a2d7e1cd7f0 · outbound

This paper cites Introducing gpt-4.1 in the api.

MagiC: Evaluating Multimodal Cognition Toward Grounded Visual Reasoning Introducing gpt-4.1 in the api

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:51:07.933231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T18:51:04.872081Z digest=sha256:cb72fe38e44479ab15fb7e4c7dd90e6ec420ff800832bfa436ef03b1bf9c1a3a

Observation 73f51922-0f6a-4e14-b238-55e7ff329eb9 · outbound

This paper cites Qvq: To see the world with wisdom, December 2024.

MagiC: Evaluating Multimodal Cognition Toward Grounded Visual Reasoning Qvq: To see the world with wisdom, December 2024

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:51:07.766881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T18:51:04.903213Z digest=sha256:3d0ba0beecc11b6b6fb57149f795c7e0a1d3d8ead5ba9e06b38775040ef1a6e0

Observation 52598951-ae8e-4590-9c5c-23e5980ed9d9 · outbound

This paper cites Uncovering the full potential of visual grounding methods in VQA.

MagiC: Evaluating Multimodal Cognition Toward Grounded Visual Reasoning Uncovering the full potential of visual grounding methods in VQA

Reference 24

Resolution
verified exact
doi, observed 2026-08-06T18:51:06.383863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T18:51:04.962971Z digest=sha256:f9346596a3bab25cf9b5552eebf7d250d8052292db194c4bc19ec4fa9d7e9e42

Observation e1e41564-69ac-402c-b641-dd779435f174 · outbound

This paper cites Visual Chain of Thought: Bridging Logical Gaps with Multimodal Infillings.

MagiC: Evaluating Multimodal Cognition Toward Grounded Visual Reasoning Visual Chain of Thought: Bridging Logical Gaps with Multimodal Infillings

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:05.057773Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:51:05.057773Z digest=sha256:cbaad11cf903c2773992adbf0c8b599db82b9f6d96127aa25f25911f010e95b5

Observation 7264deb3-fdc8-413f-8112-00ef15330c3f · outbound

This paper cites Visual CoT: Advancing Multi-Modal Language Models with a Comprehensive Dataset and Benchmark for Chain-of-Thought Reasoning.

MagiC: Evaluating Multimodal Cognition Toward Grounded Visual Reasoning Visual CoT: Advancing Multi-Modal Language Models with a Comprehensive Dataset and Benchmark for Chain-of-Thought Reasoning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:05.113854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:51:05.113854Z digest=sha256:edf7847825b3802641837196ae9966f8c032f5e217e031b725d07c839345237a

Observation 3e314ea3-e67d-44e2-9326-95c0ef7349ea · outbound

This paper cites Gemma 3 Technical Report.

MagiC: Evaluating Multimodal Cognition Toward Grounded Visual Reasoning Gemma 3 Technical Report

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:05.213221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:51:05.213221Z digest=sha256:782a7791f5f6d217418c75bb74a3877a4854c0c86fab09d81adcc7e9eb4133f2

Observation 391d53de-c277-45ab-848a-cf4e971f8c19 · outbound

This paper cites LlamaV-o1: Rethinking Step-by-step Visual Reasoning in LLMs.

MagiC: Evaluating Multimodal Cognition Toward Grounded Visual Reasoning LlamaV-o1: Rethinking Step-by-step Visual Reasoning in LLMs

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:05.257142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:51:05.257142Z digest=sha256:489b31c6a070272499256191d734d8eafb55db3d443f1ce0e59240f71c735315

Observation 51ec8827-b6fb-4c75-a226-2429fe40ba5e · outbound

This paper cites Contrastive region guidance: Improving grounding in vision-language models without training.

MagiC: Evaluating Multimodal Cognition Toward Grounded Visual Reasoning Contrastive region guidance: Improving grounding in vision-language models without training

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:05.326894Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:51:05.326894Z digest=sha256:a4bb855f864dc06d6b29bcf6c97c31af9beef8fed412c13a83daf627d657c693

Observation 2fa00a24-98fb-41ab-a6ca-c0a1990eabe7 · outbound

This paper cites Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization.

MagiC: Evaluating Multimodal Cognition Toward Grounded Visual Reasoning Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:05.402887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:51:05.402887Z digest=sha256:2541a1f3fbc9430abb4c84e869b5fadddd1fc4a2eba4a006573f048fd6d916c0

Observation 2430184c-629d-4eec-8bc3-f9736569496e · outbound

This paper cites VISCO: Benchmarking Fine-Grained Critique and Correction Towards Self-Improvement in Visual Reasoning.

MagiC: Evaluating Multimodal Cognition Toward Grounded Visual Reasoning VISCO: Benchmarking Fine-Grained Critique and Correction Towards Self-Improvement in Visual Reasoning

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-08-06T18:51:06.738821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T18:51:05.455071Z digest=sha256:7b6c6d8653b1374352441e42ee3f2a0d0177b47c99fd4de97047f52c89f3cf01

Observation 62e285fa-da03-4007-b4f9-e48be3f54dcf · outbound

This paper cites Llava-onevision-chat: Improving chat with preference learning, September 2024.

MagiC: Evaluating Multimodal Cognition Toward Grounded Visual Reasoning Llava-onevision-chat: Improving chat with preference learning, September 2024

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:51:07.629867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T18:51:05.525846Z digest=sha256:7a11e94ee0f30db8bf67b3e79310d5f6f51a5a076af1ae87082bc83db890aee8

Observation b0d75234-27f1-4d5f-b1e7-762272879e96 · outbound

This paper cites LLaVA-CoT: Let Vision Language Models Reason Step-by-Step.

MagiC: Evaluating Multimodal Cognition Toward Grounded Visual Reasoning LLaVA-CoT: Let Vision Language Models Reason Step-by-Step

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:05.654601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:51:05.654601Z digest=sha256:d2ad9dc90906c0efcbc9e65d76bf02965652989ec7b5a9c962667bd598306977

Observation 08cec980-e5df-487c-a7a2-2e261c508c7b · outbound

This paper cites Improving Visual Grounding by Encouraging Consistent Gradient-based Explanations.

MagiC: Evaluating Multimodal Cognition Toward Grounded Visual Reasoning Improving Visual Grounding by Encouraging Consistent Gradient-based Explanations

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:05.755348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:51:05.755348Z digest=sha256:5290110e0bd756b243b21a780cfb39456d1624bf4597b9782202196295ee3651

Observation a99c0d05-b26f-4792-a42a-7dc8685f882a · outbound

This paper cites Mm-vet: evaluating large multimodal models for integrated capabilities.

MagiC: Evaluating Multimodal Cognition Toward Grounded Visual Reasoning Mm-vet: evaluating large multimodal models for integrated capabilities

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:51:07.478808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T18:51:05.824802Z digest=sha256:1a66dda675084e0c1e0c2f6df84115fb595d606ad7e4d4d2adddd0831333b630

Observation 3d4c7102-8b8d-4b1d-b41b-10db977a64f9 · outbound

This paper cites Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi.

MagiC: Evaluating Multimodal Cognition Toward Grounded Visual Reasoning Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:51:07.321052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T18:51:05.946191Z digest=sha256:41f6366654879e0c17d2dff741e17dce56210e1d1840d82d77292b433f04d729

Observation ca938052-b05d-4806-b200-a1166d0a1387 · outbound

This paper cites Improve Vision Language Model Chain-of-thought Reasoning.

MagiC: Evaluating Multimodal Cognition Toward Grounded Visual Reasoning Improve Vision Language Model Chain-of-thought Reasoning

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:06.064596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:51:06.064596Z digest=sha256:aca6daa404d66a141dca1889dca1bbb36310d58c0f719a96bf79a369114ab78f

Observation 5cc0ad3c-9096-45f9-bb46-c620dd382045 · outbound

This paper cites Multimodal chain-of-thought reasoning in language models.

MagiC: Evaluating Multimodal Cognition Toward Grounded Visual Reasoning Multimodal chain-of-thought reasoning in language models

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:51:07.198965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T18:51:06.133863Z digest=sha256:69be80cbcd14771775ab48a5a04b30a37e9ed974194a8ac0df337fb2e10e24b0

Observation 6a330a3f-c247-42d8-9270-a560a91fd405 · outbound

This paper cites CoT-VLA: Visual Chain-of-Thought Reasoning for Vision-Language-Action Models.

MagiC: Evaluating Multimodal Cognition Toward Grounded Visual Reasoning CoT-VLA: Visual Chain-of-Thought Reasoning for Vision-Language-Action Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:06.202263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:51:06.202263Z digest=sha256:20ea4b5bc8f0eac46c307b9461bf223473945af22be8e81fd171cf11c0d09043

Pith citing papers

No inbound Pith citation observations are available.